Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

17 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Awesome Scientific LLM Benchmarks Awesome License: MIT

Benchmarks for evaluating large language models on scientific reasoning and discovery — across mathematics, physics & astronomy, chemistry, materials science, biology, and agentic science.

Listed alphabetically within each domain.

Contents

General / Multi-domain Science

Cross-disciplinary STEM reasoning benchmarks; a few are broad exams where science is a major subset.

Benchmark Org Year Paper Code Stars Description
AGIEval Microsoft 2023 paper code stars Human-centric benchmark from standardized exams like Gaokao, SAT, LSAT, and math competitions.
ARB DuckAI / Georgia Tech 2023 paper code stars Advanced problems across mathematics, physics, biology, chemistry, and law, including symbolic and proof items graded by rubric.
ARC (AI2 Reasoning Challenge) Allen AI (AI2) 2018 paper code stars 7,787 grade-school science exam multiple-choice questions split into Easy and Challenge sets.
C-Eval SJTU / HKUST 2023 paper code stars 13,948 Chinese multiple-choice questions across 52 disciplines and four difficulty levels.
CURIE Google 2025 paper code stars 580 expert-curated problems drawn from 429 research documents across ten long-context tasks in materials science, condensed matter physics, quantum computing, geospatial analysis, biodiversity, and proteins.
EMMA CUHK / Microsoft 2025 paper code stars 2,788 problems demanding cross-modal reasoning across mathematics, physics, chemistry, and coding for multimodal models.
FrontierScience OpenAI 2026 paper Expert-level science benchmark with an olympiad track (IPhO/IChO/IBO-level) and a rubric-graded PhD-level research track across physics, chemistry, and biology.
GAOKAO-Bench Fudan University 2023 paper code stars 2,811 Chinese college-entrance-exam questions across subjects for LLM evaluation.
GPQA NYU / Anthropic / Cohere 2023 paper code stars 448 expert-written graduate-level biology, physics, and chemistry Google-proof multiple-choice questions.
Humanity's Last Exam CAIS / Scale AI 2025 paper code stars 2,500 expert questions across 100+ subjects at the frontier of human academic knowledge.
JEEBench IIT Delhi 2023 paper code stars 515 challenging IIT-JEE Advanced physics, chemistry, and math problem-solving questions.
MMLU-Pro University of Waterloo (TIGER-Lab) 2024 paper code stars 12K reasoning-focused ten-option questions across 14 academic and STEM domains.
MMMU IN.AI / University of Waterloo / OSU 2023 paper code stars 11.5K college-level multimodal questions across six disciplines and 30 subjects.
MMSci UC Santa Barbara et al. 2024 paper code stars Figure-captioning and multiple-choice tasks built from open-access Nature Communications articles spanning 72 scientific disciplines.
OlympiadBench Tsinghua University / OpenBMB 2024 paper code stars 8,476 olympiad-level bilingual multimodal math and physics problems with expert step annotations.
OlympicArena SJTU (GAIR) 2024 paper code stars 11,163 olympiad-level problems across seven disciplines for multi-discipline cognitive reasoning.
OpenBookQA Allen AI (AI2) 2018 paper code stars Elementary science multiple-choice questions requiring core facts plus broad common-sense knowledge.
PaperMind University of Illinois Urbana-Champaign 2026 paper code stars Multimodal benchmark evaluating agent-oriented reasoning and critique over real research papers across seven domains via grounding, experimental interpretation, cross-source evidence, and critical-assessment tasks.
QASC Allen AI (AI2) 2020 paper code stars 9,980 grade-school science questions requiring retrieval and composition of two facts.
ScholarQABench University of Washington / Allen AI (AI2) 2024 paper code stars 2,967 expert-written literature-search queries with 208 long-form expert answers across computer science, physics, neuroscience, and biomedicine, scored on citation-supported synthesis.
SciArena Yale NLP / Allen AI (AI2) 2025 paper code stars Open platform where researchers vote head-to-head on model answers to literature-grounded science questions, paired with SciArena-Eval for meta-evaluating automated judges.
SciAssess DP Technology (deepmodeling) 2024 paper code stars Scientific literature analysis over biology, chemistry, materials, and medicine at three levels.
SciBench UCLA 2023 paper code stars 692 open-ended college-level chemistry, physics, and math problems requiring multi-step reasoning.
ScienceQA UCLA / Allen AI / ASU 2022 paper code stars ~21K multimodal multiple-choice science questions with lecture and chain-of-thought explanations.
SciEval Fudan University (OpenDFM) 2023 paper code stars ~18K multi-level questions testing scientific knowledge across chemistry, physics, and biology.
SciFact Allen AI (AI2) 2020 paper code stars 1,409 expert-written scientific claims verified against research abstracts, with rationales.
SciKnowEval Zhejiang University (HICAI) 2024 paper code stars Tens of thousands of problems across five cognitive levels in biology, chemistry, physics, and materials.
SciQ Allen AI (AI2) 2017 paper 13,679 crowdsourced science multiple-choice questions across physics, chemistry, and biology with evidence.
SDE Cornell / Princeton / Stanford / MIT / Toronto 2025 paper code stars Scenario-grounded scientific-discovery benchmark of 43 scenarios and 1,125 questions across biology, chemistry, materials, and physics, plus 8 project-level hypothesis, experiment-design, and interpretation tasks.
SuperGPQA ByteDance Seed / M-A-P 2025 paper code stars 26,529 graduate-level questions spanning 285 disciplines, including under-evaluated long-tail fields.
TheoremQA University of Waterloo (TIGER-Lab) 2023 paper code stars 800 questions applying 350+ theorems across math, physics, EE/CS, and finance.
Xiezhi Fudan University 2023 paper code stars 249,587 questions spanning 516 disciplines across 13 categories, continuously updated.

Back to top

Mathematics

Arithmetic, competition, olympiad, and frontier / formal-proof mathematics.

Benchmark Org Year Paper Code Stars Description
CHAMP MIT 2024 paper code stars 270 high-school competition problems annotated with concepts and problem-specific hints.
CombiBench Moonshot AI / Numina 2025 paper code stars 100 competition combinatorics problems formalized in Lean4, a domain underrepresented by existing formal benchmarks.
FIMO Peking University / Huawei 2023 paper code stars 149 IMO-shortlist problems formalized in Lean with informal statements for olympiad-level theorem proving.
FormalMATH SphereLab / M-A-P 2025 paper code stars 5,560 formally verified Lean4 statements from olympiad to undergraduate across algebra, calculus, and number theory.
FrontierMath Epoch AI 2024 paper Hundreds of unpublished expert-crafted research-level math problems resistant to guessing.
GSM-Symbolic Apple 2024 paper code stars Symbolic templates over GSM8K that regenerate each problem with new names, values, and extra clauses to test whether grade-school math accuracy survives superficial perturbation.
GSM8K OpenAI 2021 paper code stars 8.5K grade-school arithmetic word problems requiring multi-step reasoning.
HARDMath Harvard 2024 paper code stars Graduate applied-math problems requiring asymptotic and approximation methods where leading LLMs score below 45%.
HARP UCL 2024 paper code stars US competition math problems (AMC/AIME/USAMO) with human-written ground-truth solutions.
IMO-Bench Google DeepMind 2025 paper code stars Olympiad-level suite testing final answers, proof-writing, grading, and Lean formal proofs, vetted by IMO medalists.
Lean Workbook Shanghai AI Lab 2024 paper About 57K formal-informal Lean4 problem pairs auto-formalized from competition math.
LeanDojo Caltech / NVIDIA 2023 paper code stars 98,734 theorems and proofs from Lean mathlib with premise annotations for retrieval-augmented theorem proving.
MATH UC Berkeley 2021 paper code stars 12,500 competition mathematics problems (AMC/AIME level) with step-by-step solutions.
MATH-500 OpenAI 2023 paper code stars 500-problem held-out subset of MATH widely used for LLM evaluation.
MATH-Vision (MATH-V) CUHK MMLab / Shanghai AI Lab 2024 paper code stars 3,040 real competition problems with visual context, spanning 16 mathematical disciplines and five difficulty levels.
MathArena ETH Zurich / INSAIT 2025 paper code stars Evaluates models only on competitions held after their release date to rule out contamination, covering 162 problems from seven competitions plus human-graded IMO proof writing.
MathBench Shanghai AI Laboratory 2024 paper code stars Hierarchical bilingual benchmark spanning arithmetic to college-level theory and application.
MATHCHECK XJTLU / HKUST 2024 paper code stars Checklist benchmark testing task generalization and reasoning robustness beyond end-to-end answer accuracy.
MathOdyssey NetMind.AI 2024 paper code stars 387 expert-crafted high-school, university, and olympiad-level math problems.
MathQA UW / Allen AI 2019 paper code stars 37K multiple-choice math word problems annotated with executable operation programs.
MathVerse CUHK MMLab / Shanghai AI Lab 2024 paper code stars 2,612 visual math problems in six diagram/text variants probing whether MLLMs truly interpret figures.
MathVista UCLA / University of Washington / Microsoft 2023 paper code stars Mathematical reasoning benchmark combining visual contexts (figures, charts, diagrams) with problems.
miniF2F OpenAI 2021 paper code stars 488 olympiad-level (AMC/AIME/IMO) problems formalized across multiple proof assistants.
Omni-MATH Peking University 2024 paper code stars 4,428 olympiad-level competition problems across 33+ subdomains and difficulty tiers.
ProofNet Yale University 2023 paper code stars 371 undergraduate theorems for autoformalization and formal proving in Lean.
Putnam-AXIOM Stanford 2025 paper Putnam competition problems plus programmatically perturbed variations giving contamination-resistant unseen instances.
PutnamBench UT Austin 2024 paper code stars 1,600+ Putnam competition problems formalized in Lean, Isabelle, and Coq.
We-Math BUPT / Tencent 2024 paper code stars 6.5K visual math problems decomposed into 67 hierarchical knowledge concepts to diagnose reasoning versus memorization.

Back to top

Physics & Astronomy

Physics olympiad, graduate physics, computational physics, and astronomy.

Benchmark Org Year Paper Code Stars Description
ABench-Physics Zhejiang University / Ant Group 2025 paper 500 hard static plus dynamic-variant physics problems testing reasoning and generalization robustness.
Astro-QA ACMIS Lab 2025 paper code stars About 2,700 bilingual astronomy questions across six types spanning astrophysics, celestial mechanics, and astrometry.
AstroMLab 1 AstroMLab Collaboration 2024 paper 4,425 astronomy multiple-choice questions from Annual Reviews testing LLM astronomical knowledge.
AstroVisBench NSF-Simons CosmicAI Institute 2025 paper code stars Evaluates LLMs on end-to-end astronomy computing workflows and scientific result visualization.
AtmosSci-Bench HKUST 2025 paper code stars Atmospheric-science problems spanning dynamics, physics, hydrology, geophysics, and oceanography via templated questions.
ClimaQA UC San Diego 2024 paper code stars Graduate-level climate science QA generated from textbooks with climate scientists in the loop: 566 expert-validated Gold questions plus 3,000 synthetic Silver questions in MCQ, freeform, and cloze form.
CMPhysBench Chinese Academy of Sciences (IOP) 2025 paper code stars 520+ graduate condensed-matter physics problems with a partial-credit metric; top models score below 30%.
CritPt Argonne National Laboratory / UIUC (50+ physicists, 30+ institutions) 2025 paper code stars 71 composite research-project challenges and 190 modular checkpoints authored by 50+ active physicists, auto-graded via numerical, symbolic, and code evaluations.
DiscoverPhysics Princeton / NYU / Flatiron Institute / Polymathic AI 2026 paper code stars 22 simulated worlds with non-standard physics where an agent proposes initial conditions, observes noisy N-body trajectories, and submits a natural-language law plus a Python implementation.
FEABench Google Research / Harvard 2025 paper code stars Finite-element analysis problems solved end-to-end by driving the COMSOL Multiphysics API: 15 manually verified problems (Gold) plus 200 algorithmically parsed tasks (Large).
Gravity-Bench-v1 University of Toronto 2025 paper code stars Agents plan budgeted observations of simulated binary-star systems to discover the underlying gravitational physics.
HiPhO CUHK / Shanghai AI Lab 2025 paper code stars 13 recent physics olympiad exams with step-level grading and human medal-based comparison for (M)LLMs.
LLM-SRBench Virginia Tech / CMU 2025 paper code stars 239 problems across four science domains testing genuine LLM equation discovery over memorized formulas.
NewtonBench HKUST 2025 paper code stars 324 tasks across twelve physics domains testing LLM agents discovering scientific laws via interactive experimentation.
PHYBench Peking University 2025 paper code stars 500 original physics problems from high-school to olympiad level with expression-edit-distance scoring.
PhysGym KAUST 2025 paper code stars 97 simulated interactive physics-discovery problems with four controlled levels of prior knowledge.
PHYSICS Yale University 2025 paper code stars 1,297 expert-annotated university-level physics problems across six core areas with automated evaluation.
PhysReason Xi'an Jiaotong University 2025 paper code stars 1,200 physics problems with step-level automatic scoring for multi-step physics reasoning.
PRL-Bench Shanghai Jiao Tong University 2026 paper End-to-end physics research benchmark built from ~100 recent Physical Review Letters papers across five subfields, each turned into a long-horizon task scored by an LLM-as-judge.
TPBench University of Wisconsin-Madison 2025 paper 57 theoretical physics problems in high-energy theory and cosmology, undergraduate to research level.
UGPhysics HKUST 2025 paper code stars 5,520 bilingual undergraduate physics problems across 13 subjects with rule-based judging.

Back to top

Chemistry

Molecular property, reaction, retrosynthesis, safety, and chemical knowledge.

Benchmark Org Year Paper Code Stars Description
ChemBench LAMALab, University of Jena (Jablonka group) 2024 paper code stars 2,700+ curated questions probing LLM chemical knowledge and reasoning against expert chemists.
ChemCoTBench IDEA Research / PKU / CUHK 2025 paper code stars Step-wise chemical reasoning across molecular understanding, editing, optimization, and reaction prediction.
ChemEval USTC 2024 paper code stars Multi-level chemistry benchmark spanning 42 tasks across four progressive difficulty levels for LLMs.
ChemLLMBench University of Notre Dame et al. 2023 paper code stars Eight chemistry tasks including reaction prediction, retrosynthesis, and molecule captioning for LLMs.
ChemSafetyBench Peking University 2024 paper Safety benchmark testing LLM handling of hazardous-chemical property, legality, and synthesis queries.
MaCBench LAMALab (Jena) / IIT Delhi 2024 paper code stars 1,100+ image-question pairs probing multimodal LLM limitations across chemistry and materials workflows.
MolCap-Arena Genentech / UIUC 2024 paper code stars Battle-based benchmark scoring 20+ LLMs on molecule captions augmenting molecular property prediction.
MoleculeQA IDEA Research 2024 paper code stars 62K QA pairs over 23K molecules evaluating factual accuracy of LLM molecular comprehension.
MolTextQA UIUC 2024 paper code stars 500K QA pairs over 240K PubChem molecules evaluating molecule-structure-to-text understanding and retrieval.
MOOSE-Chem NTU Singapore / Shanghai AI Lab 2024 paper code stars Tests whether LLMs rediscover unseen chemistry hypotheses from 51 Nature/Science-level papers given background information.
ScholarChemQA KAUST / Notre Dame 2024 paper code stars 40K abstract-derived chemistry research questions exposing LLM limits in comprehending scholarly literature.
SMolInstruct / LlaSMol OSU NLP Group 2024 paper code stars Large-scale instruction dataset of 14 small-molecule chemistry tasks used to train and evaluate LlaSMol.
TOMG-Bench HK PolyU / Shanghai AI Lab 2024 paper code stars Evaluates LLMs on text-based open-domain molecule generation across editing, optimization, and customized generation.

Back to top

Materials Science

Crystals, materials property prediction, and materials-science knowledge.

Benchmark Org Year Paper Code Stars Description
ALDbench Argonne National Laboratory 2024 paper code stars Expert open-ended QA benchmark evaluating LLMs on atomic-layer-deposition synthesis for accuracy and specificity.
AtomWorld USTC / Shanghai AI Laboratory / UNSW 2025 paper code stars Evaluates LLM spatial reasoning on crystalline materials via ten atomic-structure editing actions across four modeling categories over CIF files with verifiable metrics.
LLM4Mat-Bench Princeton (Vertaix) 2024 paper code stars Largest benchmark for LLM materials property prediction over 1.9M crystals and 45 properties.
MaScQA IIT Delhi (M3RG) 2024 paper code stars 650 undergraduate-level materials science questions across 14 categories for evaluating LLM knowledge.
MatCha CUHK-Shenzhen 2025 paper code stars 1,500 expert questions across 21 tasks for materials characterization image understanding by multimodal models.
MatSci-NLP Mila / Universite de Montreal 2023 paper code stars Seven materials-science NLP tasks evaluating language models via a unified text-to-schema approach.
MatSciBench UCLA 2025 paper code stars College-level materials-science reasoning benchmark of 1,340 problems spanning six fields, including multimodal questions.
MatText LamaLab (Jena) / Intel Labs 2024 paper code stars Benchmarking framework evaluating language models on crystal property prediction across nine text representations.
MatViX Duke University 2024 paper code stars Benchmarks vision-language models on multimodal information extraction from visually-rich materials-science articles.
MatVQA Mila / Universite de Montreal 2025 paper 1,325 questions testing multimodal models on materials imagery like microscopy and diffraction with multi-step reasoning.

Back to top

Biology & Life Sciences

Genomics, proteins, bioinformatics agents, protocols, and research biology.

Benchmark Org Year Paper Code Stars Description
BioCoder Yale University (Gerstein Lab) 2023 paper code stars Fuzz-tested benchmark evaluating LLMs on generating bioinformatics code with cross-file dependencies and domain knowledge.
BioLLMBench UCLA 2023 paper 2,160 runs evaluating GPT-4, Bard, and LLaMA across 24 bioinformatics tasks and metrics.
BioMaze Peking University 2025 paper code stars 5.1K real-research pathway problems testing LLM reasoning over biological pathways and perturbations.
BioML-bench ScienceMachine 2025 paper code stars AI agents build end-to-end biomedical ML pipelines across protein, omics, imaging, and drug tasks.
Biomni Stanford 2025 paper code stars General biomedical agent with Biomni-Eval: 433 instances across 10 reasoning tasks.
BioMysteryBench Anthropic 2026 paper 99 expert-written bioinformatics tasks over raw datasets, judged on the final biological conclusion, not the path.
BioPlanner FutureHouse / Francis Crick / Oxford 2023 paper code stars BioProt dataset automatically evaluating LLMs on generating biology experimental protocols as executable pseudocode.
BioProBench Peking University 2025 paper code stars Tests LLMs on biological-protocol QA, step ordering, error correction, generation, and reasoning across 556K instances.
BixBench FutureHouse 2025 paper code stars Bioinformatics agents tackle 53 real analysis scenarios with ~300 open-answer research questions.
CellVoyager Stanford 2025 paper code stars CellBench: 76 scRNA-seq studies test agents predicting which analyses the authors performed.
GeneBench-Pro OpenAI 2026 paper 129 messy, judgment-intensive computational-biology problems across ten genomics domains on realistic datasets.
GeneGPT NCBI 2023 paper code stars Teaches LLMs to call NCBI Web APIs; adds GeneHop, evaluated on GeneTuring.
GeneTuring Columbia University 2023 paper code stars 16 genomics tasks, 1,600 questions probing LLM genomic knowledge and reasoning.
Genome-Bench Princeton / Stanford 2025 paper 3,000+ genome-engineering multiple-choice questions mined from expert forum discussions for evaluating LLM reasoning.
GenoTEX UIUC 2024 paper code stars Expert-curated gene-expression tasks: dataset selection, preprocessing, and gene-trait statistical analysis.
HealthBench OpenAI 2025 paper 5,000 multi-turn physician-crafted health conversations scored by rubric criteria for LLM performance and safety.
LAB-Bench FutureHouse 2024 paper code stars 2,400+ MCQs across literature, figures, databases, protocols, and DNA/protein sequence tasks.
LABBench2 FutureHouse (Edison Scientific) 2026 paper code stars Successor to LAB-Bench with ~1,900 biology-research tasks in realistic contexts (literature, figures, protocols, databases), giving a sharp difficulty jump over LAB-Bench.
LifeSciBench OpenAI 2026 paper 750 expert-authored free-response life-science research tasks across seven biological domains, rubric-graded.
ProteinLMBench Shanghai Jiao Tong University 2024 paper code stars 944 verified MCQs assessing LLM comprehension of protein sequences and descriptions.
VirBench Anthropic (with NCBI, Broad Institute, Pachter Lab) 2026 paper 120 curated viral-sequence retrieval queries across ~40 pathogens testing whether LLM agents pull correct ground-truth sequence data from NCBI Virus.

Back to top

Agentic Science & AI Research

LLM agents that write research code, run data analyses, attempt autonomous discovery, and conduct ML/AI research.

Benchmark Org Year Paper Code Stars Description
AAAR-1.0 Penn State et al. 2024 paper code stars Expert-annotated tasks assessing LLMs on equation inference, experiment design, and paper-weakness identification.
AstaBench Allen AI (AI2) 2025 paper code stars 2,400+ problems across eleven benchmarks evaluating agents over the full scientific research pipeline.
BLADE University of Washington 2024 paper code stars Twelve datasets and research questions evaluating agents' analytical decisions against expert ground truth.
CORE-Bench Princeton 2024 paper code stars 270 tasks from 90 papers testing agents on computationally reproducing published scientific results.
Curie University of Michigan 2025 paper code stars Framework and 46-question benchmark for rigorous, automated scientific experimentation across four CS domains.
DiscoveryBench Allen AI (AI2) 2024 paper code stars 264 real plus 903 synthetic tasks for data-driven hypothesis discovery across six domains.
DiscoveryWorld Allen AI (AI2) 2024 paper code stars Simulated environment with 120 tasks requiring full cycles of hypothesis, experiment, and analysis.
DSBench UT Dallas / Tencent AI Lab 2024 paper code stars 540 realistic data-analysis and data-modeling tasks sourced from ModelOff and Kaggle competitions.
EXP-Bench University of Michigan 2025 paper Benchmarks agents on designing, implementing, and analyzing end-to-end AI research experiments from publications.
HeurekaBench EPFL (MLBio Lab) 2026 paper code stars Builds benchmarks of open-ended research questions grounded in real studies and their code to evaluate end-to-end AI co-scientist agents (instantiated as sc-HeurekaBench in single-cell biology).
InnovatorBench GAIR-NLP (SJTU) 2025 paper code stars Evaluates agents on end-to-end innovative LLM-research tasks like data construction, loss design, and reward design.
LiveIdeaBench Renmin University of China 2024 paper code stars Benchmarks scientific idea generation from minimal keyword context across five dimensions: originality, feasibility, fluency, flexibility, and clarity.
LMR-Bench UT Dallas 2025 paper code stars 28 tasks from 23 NLP papers testing whether agents reproduce masked research code, verified by unit tests.
MLAgentBench Stanford 2023 paper code stars 13 ML experimentation tasks where agents read, write, and execute code to improve performance.
MLE-bench OpenAI 2024 paper code stars 75 Kaggle ML-engineering competitions testing agents against human leaderboards.
MLGym Meta 2025 paper code stars Gym framework with 13 open-ended AI research tasks spanning vision, NLP, RL, and game theory.
MLR-Bench National University of Singapore 2025 paper code stars 201 workshop-derived tasks evaluating agents across idea generation, experimentation, and paper writing.
MLRC-Bench University of Michigan 2025 paper code stars Benchmarks agents on proposing and coding novel methods for seven recent ML research competition problems.
Paper2Code KAIST / DeepAuto.ai 2025 paper code stars Multi-agent framework and benchmark generating functional code repositories that reproduce ML research papers.
PaperBench OpenAI 2025 paper Agents replicate 20 ICML 2024 papers from scratch, graded on 8,316 rubric criteria.
RE-Bench METR 2024 paper code stars Seven open-ended ML research-engineering environments comparing agents against 61 human experts.
ResearchBench Shanghai AI Lab / NTU 2025 paper Benchmarks scientific discovery via inspiration retrieval, hypothesis composition, and ranking across twelve disciplines.
ResearchCodeBench Stanford University 2025 paper code stars Challenges LLMs to implement novel contributions from recent ML papers by completing TODO code snippets.
RExBench Boston University 2025 paper code stars Twelve tasks where agents autonomously implement novel research extensions to existing papers' codebases.
SciAgentGym Fudan NLP Group 2026 paper code stars Agentic science benchmark pairing an environment of 1,780 domain-specific tools across natural-science disciplines with a tiered suite from elementary tool actions to long-horizon workflows.
SciCode UIUC / Princeton / Argonne National Lab 2024 paper code stars 80 real research coding problems across 16 natural-science subfields, split into 338 subproblems.
ScienceAgentBench Ohio State (OSU-NLP) 2024 paper code stars 102 data-driven discovery tasks from 44 peer-reviewed papers, evaluating agents that write Python programs.
ScienceWorld University of Arizona / Microsoft Research / Allen AI (AI2) 2022 paper code stars Interactive text environment with 30 elementary-science tasks where agents must run the experiment — melt a substance, test conductivity, breed a plant — instead of reciting the answer.
Scientist-Bench University of Hong Kong 2025 paper code stars Evaluates fully autonomous idea-to-paper research systems across CV, NLP, data mining, and IR against expert papers.
SciReplicate-Bench King's College London 2025 paper code stars Benchmarks agents on reproducing executable code for algorithms described in 36 recent research papers.
SUPER Allen AI (AI2) 2024 paper code stars Evaluates agents on setting up and executing tasks from research code repositories.

Back to top

Contributing

Found a missing benchmark or an error? Edit only data/benchmarks.yaml and open a PR — the README, website, and examples regenerate automatically. See CONTRIBUTING.md for the schema and inclusion bar.

License

Released under the MIT License. Each entry links to its original paper and code; benchmark metadata is drawn from public sources.

About

A curated, accuracy-first list of benchmarks for evaluating LLMs on scientific reasoning and discovery — math, physics, chemistry, materials, biology, and agentic science.

Topics

Resources

Contributing

Stars

31 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages