Drug Discovery Benchmark | Insilico Medicine

Benchmarking AI for Drug Discovery

A unified benchmarking platform combining curated public datasets with Insilico Medicine's proprietary benchmarks to comprehensively evaluate AI across the drug discovery pipeline. A research initiative by Insilico Medicine.

Hosted Benchmarks

Drug Discovery Benchmark Benchmarks Explorer

Interactive exploration of benchmarks across Biology, Chemistry, Materials, and Longevity. Browse and compare model performance with built-in visualizations.

Global Leaderboard

Models Ranked

Top Model

Included benchmarks (229)

Benchmark Metric Range
Aging Prediction Balanced accuracy [0, 1]
GTEx Expression accuracy_score [0, 1]
GTEx Pairwise Binary Balanced accuracy [0, 1]
GTEx Pairwise Ternary accuracy_score [0, 1]
Longevity Synergy (Full) Balanced accuracy [0, 1]
Longevity Synergy (Minimal) Balanced accuracy [0, 1]
Methylation Age Choice accuracy_score [0, 1]
Methylation Age Pairwise Balanced accuracy [0, 1]
Methylation Age Regression Spearman correlation [-1, 1]
NHANES Mortality Classification Balanced accuracy [0, 1]
NHANES Pairwise Comparison Balanced accuracy [0, 1]
NHANES Time-to-Event accuracy_score [0, 1]
NHANES TTE Regression Spearman correlation [-1, 1]
Olink Pairwise Comparison Balanced accuracy [0, 1]
Olink Protein Classification Balanced accuracy [0, 1]
Synergy Regression Spearman correlation [-1, 1]
TCGA Survival Prediction Balanced accuracy [0, 1]
Target Prediction for Cancer Diseases Precision@K [0, 1]
... ... ...

Ranked Table

# Model Insilico Score Disease Biology Chemical Synthesis Medicinal Chemistry Clinical Trials Biologics
1 GPT 5.5 69.9 34 30 72 81 56
2 Gemini 3.1 Pro 63.7 32 30 67 78 54
3 Claude Opus 4.8 62.4 28 26 63 78 55
4 Grok 4.5 60.7 29 28 61 78 52
5 Grok 4.3 56.2 16 18 58 71 53
6 GPT 5.6 Sol 28.4 10 70 54
7 Kimi K3 24.7 63

About Drug Discovery Benchmark

Drug Discovery Benchmark brings together the best of both worlds — established public benchmarks and Insilico Medicine's proprietary evaluations — to provide the most comprehensive assessment of AI in drug discovery.

By combining publicly available datasets with proprietary tasks grounded in real-world drug development, this platform offers a uniquely balanced view of model strengths across every stage of the pipeline.

Hybrid Benchmarks

A curated mix of public community benchmarks and proprietary Insilico evaluations.

LLM Evaluation

Comprehensive testing of large language models on domain-specific tasks.

Transparent Metrics

Clear, reproducible evaluation metrics with public leaderboards.

Pipeline Coverage

Benchmarks spanning target discovery through clinical development, covering every key stage.