Insilico Bench | Insilico Medicine
Insilico's Proprietary AI Benchmarks
Proprietary benchmarks developed by Insilico Medicine to evaluate AI systems across the drug discovery pipeline — from target identification to clinical development. A research initiative by Insilico Medicine.
Insilico Bench Benchmarks Explorer
Interactive exploration of benchmarks across Biology, Chemistry, Materials, and Longevity. Browse and compare model performance with built-in visualizations.
Models Ranked
16 of 17 total Models Ranked
198 of 203 total Benchmarks
5 active Categories
Top Model: GPT 5.5
Insilico score: 66.4
Included benchmarks (198)
- Aging Prediction: Balanced accuracy [0, 1]
- GTEx Expression: accuracy_score [0, 1]
- GTEx Pairwise Binary: Balanced accuracy [0, 1]
- GTEx Pairwise Ternary: accuracy_score [0, 1]
- Longevity Synergy (Full): Balanced accuracy [0, 1]
- Longevity Synergy (Minimal): Balanced accuracy [0, 1]
- Methylation Age Choice: accuracy_score [0, 1]
- Methylation Age Pairwise: Balanced accuracy [0, 1]
- Methylation Age Regression: Spearman correlation [-1, 1]
- NHANES Mortality Classification: Balanced accuracy [0, 1]
- NHANES Pairwise Comparison: Balanced accuracy [0, 1]
- NHANES Time-to-Event: accuracy_score [0, 1]
- NHANES TTE Regression: Spearman correlation [-1, 1]
- Olink Pairwise Comparison: Balanced accuracy [0, 1]
- Olink Protein Classification: Balanced accuracy [0, 1]
- TCGA Survival Prediction: Balanced accuracy [0, 1]
- ...
Ranked Table
| # | Model | Insilico Score | Biology | Affinity and Binding | Chemical Synthesis | ADMET, PK & Safety | Clinical Trials |
|---|---|---|---|---|---|---|---|
| 1 | GPT 5.5 | 66.4 | 35 | 71 | 30 | 70 | 81 |
| 2 | Claude Opus 4.8 | 58.6 | 30 | 57 | 26 | 69 | 78 |
| 3 | Claude Opus 4.6 | 58.5 | 59 | 62 | 9 | 63 | 72 |
| 4 | Claude Opus 4.5 | 57.9 | 63 | 59 | 9 | 66 | 73 |
| 5 | Claude Opus 4.7 | 56.9 | — | 70 | 12 | 70 | — |
| 6 | DeepSeek v3.2 | 56.3 | 56 | 62 | 1 | 59 | 68 |
| 7 | Claude Sonnet 4.5 | 55.9 | 61 | 62 | 6 | 62 | 59 |
| 8 | GPT 5.2 | 55.1 | 57 | 62 | 4 | 61 | 53 |
| 9 | Grok 4.1 | 55.0 | 51 | 61 | 14 | 59 | 60 |
| 10 | Kimi K2.5 | 52.8 | 58 | 56 | 9 | 61 | 64 |
| ... | ... | ... | ... | ... | ... | ... | ... |
Insilico's Proprietary AI Evaluation Suite
Insilico Bench is Insilico Medicine's proprietary benchmarking platform, built to rigorously evaluate AI models on real-world drug discovery tasks using internal datasets and domain expertise.
Proprietary Data
Benchmarks built on Insilico's internal drug discovery datasets and pipelines.
LLM Evaluation
Comprehensive testing of large language models on domain-specific tasks.
Transparent Metrics
Clear, reproducible evaluation metrics with public leaderboards.
Industry-Proven
Grounded in Insilico Medicine's track record of AI-driven drug candidates in clinical trials.