Insilico Science AI Bench | Insilico Medicine
Advancing AI Across the Sciences
Evaluating AI in Science & Discovery
A comprehensive benchmarking platform for AI in science — spanning drug discovery, biology, chemistry, materials science, and beyond. A research initiative by Insilico Medicine.
Hosted Benchmarks
Insilico Science AI Bench Benchmarks Explorer
Interactive exploration of benchmarks across Biology, Chemistry, Materials, and Longevity. Browse and compare model performance with built-in visualizations.
Global Leaderboard
- Models Ranked: 8 of 12 total
- Benchmarks: 253 of 296 total
- Categories: 9 active
Top Model: GPT 5.5
Insilico Score: 67.6
Leaderboard Settings
- Scoring Mode: Insilico Score, Relative Insilico Score, Mean Rank
284 benchmarks included in computing the Insilico Score.
Included benchmarks (284)
- Aging Prediction: Balanced accuracy[0, 1]
- GTEx Expression: accuracy_score[0, 1]
- NHANES Mortality Classification: Balanced accuracy[0, 1]
- TCGA Survival Prediction: Balanced accuracy[0, 1]
- ...
Ranked Table
| # | Model | Insilico Score | Biology | Affinity and Binding | Chemical Synthesis | ADMET, PK & Safety | Clinical Trials | Biologics | Materials | Longevity & Aging | Molecular Design & Optimization |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT 5.5 | 67.6 | 34 | 71 | 29 | 72 | 81 | 56 | 88 | 66 | 49 |
| 2 | Gemini 3.1 Pro | 60.3 | 32 | 64 | 29 | 73 | 78 | 54 | — | 62 | 50 |
| 3 | Claude Opus 4.8 | 59.1 | 28 | 58 | 25 | 70 | 78 | 55 | — | 61 | 49 |
| 4 | Grok 4.5 | 58.1 | 29 | 56 | 28 | 66 | 78 | 52 | — | 68 | 47 |
| 5 | Claude Opus 4.7 | 56.4 | — | 71 | 16 | 70 | — | 56 | 84 | — | 33 |
| ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... |
About Insilico Science AI Bench
Insilico Science AI Bench is a broad benchmarking platform developed by Insilico Medicine to evaluate AI across scientific domains — from drug discovery and biology to materials science and longevity research.
- Multi-Domain Science: Benchmarks spanning drug discovery, chemistry, biology, materials science, and longevity.
- AI Model Evaluation: Testing large language models and specialized AI systems on diverse scientific tasks.
- Transparent Metrics: Clear, reproducible evaluation metrics with public leaderboards.
- Research-Backed: All benchmarks supported by peer-reviewed scientific publications.