Benchmarks for evaluating large language models on scientific reasoning and discovery — across mathematics, physics & astronomy, chemistry, materials science, biology, and agentic science.
Listed alphabetically within each domain.
- General / Multi-domain Science
- Mathematics
- Physics & Astronomy
- Chemistry
- Materials Science
- Biology & Life Sciences
- Agentic Science & AI Research
Cross-disciplinary STEM reasoning benchmarks; a few are broad exams where science is a major subset.
| Benchmark | Org | Year | Paper | Code | Stars | Description |
|---|---|---|---|---|---|---|
| AGIEval | Microsoft | 2023 | paper | code | Human-centric benchmark from standardized exams like Gaokao, SAT, LSAT, and math competitions. | |
| ARB | DuckAI / Georgia Tech | 2023 | paper | code | Advanced problems across mathematics, physics, biology, chemistry, and law, including symbolic and proof items graded by rubric. | |
| ARC (AI2 Reasoning Challenge) | Allen AI (AI2) | 2018 | paper | code | 7,787 grade-school science exam multiple-choice questions split into Easy and Challenge sets. | |
| C-Eval | SJTU / HKUST | 2023 | paper | code | 13,948 Chinese multiple-choice questions across 52 disciplines and four difficulty levels. | |
| CURIE | 2025 | paper | code | 580 expert-curated problems drawn from 429 research documents across ten long-context tasks in materials science, condensed matter physics, quantum computing, geospatial analysis, biodiversity, and proteins. | ||
| EMMA | CUHK / Microsoft | 2025 | paper | code | 2,788 problems demanding cross-modal reasoning across mathematics, physics, chemistry, and coding for multimodal models. | |
| FrontierScience | OpenAI | 2026 | paper | — | — | Expert-level science benchmark with an olympiad track (IPhO/IChO/IBO-level) and a rubric-graded PhD-level research track across physics, chemistry, and biology. |
| GAOKAO-Bench | Fudan University | 2023 | paper | code | 2,811 Chinese college-entrance-exam questions across subjects for LLM evaluation. | |
| GPQA | NYU / Anthropic / Cohere | 2023 | paper | code | 448 expert-written graduate-level biology, physics, and chemistry Google-proof multiple-choice questions. | |
| Humanity's Last Exam | CAIS / Scale AI | 2025 | paper | code | 2,500 expert questions across 100+ subjects at the frontier of human academic knowledge. | |
| JEEBench | IIT Delhi | 2023 | paper | code | 515 challenging IIT-JEE Advanced physics, chemistry, and math problem-solving questions. | |
| MMLU-Pro | University of Waterloo (TIGER-Lab) | 2024 | paper | code | 12K reasoning-focused ten-option questions across 14 academic and STEM domains. | |
| MMMU | IN.AI / University of Waterloo / OSU | 2023 | paper | code | 11.5K college-level multimodal questions across six disciplines and 30 subjects. | |
| MMSci | UC Santa Barbara et al. | 2024 | paper | code | Figure-captioning and multiple-choice tasks built from open-access Nature Communications articles spanning 72 scientific disciplines. | |
| OlympiadBench | Tsinghua University / OpenBMB | 2024 | paper | code | 8,476 olympiad-level bilingual multimodal math and physics problems with expert step annotations. | |
| OlympicArena | SJTU (GAIR) | 2024 | paper | code | 11,163 olympiad-level problems across seven disciplines for multi-discipline cognitive reasoning. | |
| OpenBookQA | Allen AI (AI2) | 2018 | paper | code | Elementary science multiple-choice questions requiring core facts plus broad common-sense knowledge. | |
| PaperMind | University of Illinois Urbana-Champaign | 2026 | paper | code | Multimodal benchmark evaluating agent-oriented reasoning and critique over real research papers across seven domains via grounding, experimental interpretation, cross-source evidence, and critical-assessment tasks. | |
| QASC | Allen AI (AI2) | 2020 | paper | code | 9,980 grade-school science questions requiring retrieval and composition of two facts. | |
| ScholarQABench | University of Washington / Allen AI (AI2) | 2024 | paper | code | 2,967 expert-written literature-search queries with 208 long-form expert answers across computer science, physics, neuroscience, and biomedicine, scored on citation-supported synthesis. | |
| SciArena | Yale NLP / Allen AI (AI2) | 2025 | paper | code | Open platform where researchers vote head-to-head on model answers to literature-grounded science questions, paired with SciArena-Eval for meta-evaluating automated judges. | |
| SciAssess | DP Technology (deepmodeling) | 2024 | paper | code | Scientific literature analysis over biology, chemistry, materials, and medicine at three levels. | |
| SciBench | UCLA | 2023 | paper | code | 692 open-ended college-level chemistry, physics, and math problems requiring multi-step reasoning. | |
| ScienceQA | UCLA / Allen AI / ASU | 2022 | paper | code | ~21K multimodal multiple-choice science questions with lecture and chain-of-thought explanations. | |
| SciEval | Fudan University (OpenDFM) | 2023 | paper | code | ~18K multi-level questions testing scientific knowledge across chemistry, physics, and biology. | |
| SciFact | Allen AI (AI2) | 2020 | paper | code | 1,409 expert-written scientific claims verified against research abstracts, with rationales. | |
| SciKnowEval | Zhejiang University (HICAI) | 2024 | paper | code | Tens of thousands of problems across five cognitive levels in biology, chemistry, physics, and materials. | |
| SciQ | Allen AI (AI2) | 2017 | paper | — | — | 13,679 crowdsourced science multiple-choice questions across physics, chemistry, and biology with evidence. |
| SDE | Cornell / Princeton / Stanford / MIT / Toronto | 2025 | paper | code | Scenario-grounded scientific-discovery benchmark of 43 scenarios and 1,125 questions across biology, chemistry, materials, and physics, plus 8 project-level hypothesis, experiment-design, and interpretation tasks. | |
| SuperGPQA | ByteDance Seed / M-A-P | 2025 | paper | code | 26,529 graduate-level questions spanning 285 disciplines, including under-evaluated long-tail fields. | |
| TheoremQA | University of Waterloo (TIGER-Lab) | 2023 | paper | code | 800 questions applying 350+ theorems across math, physics, EE/CS, and finance. | |
| Xiezhi | Fudan University | 2023 | paper | code | 249,587 questions spanning 516 disciplines across 13 categories, continuously updated. |
Arithmetic, competition, olympiad, and frontier / formal-proof mathematics.
| Benchmark | Org | Year | Paper | Code | Stars | Description |
|---|---|---|---|---|---|---|
| CHAMP | MIT | 2024 | paper | code | 270 high-school competition problems annotated with concepts and problem-specific hints. | |
| CombiBench | Moonshot AI / Numina | 2025 | paper | code | 100 competition combinatorics problems formalized in Lean4, a domain underrepresented by existing formal benchmarks. | |
| FIMO | Peking University / Huawei | 2023 | paper | code | 149 IMO-shortlist problems formalized in Lean with informal statements for olympiad-level theorem proving. | |
| FormalMATH | SphereLab / M-A-P | 2025 | paper | code | 5,560 formally verified Lean4 statements from olympiad to undergraduate across algebra, calculus, and number theory. | |
| FrontierMath | Epoch AI | 2024 | paper | — | — | Hundreds of unpublished expert-crafted research-level math problems resistant to guessing. |
| GSM-Symbolic | Apple | 2024 | paper | code | Symbolic templates over GSM8K that regenerate each problem with new names, values, and extra clauses to test whether grade-school math accuracy survives superficial perturbation. | |
| GSM8K | OpenAI | 2021 | paper | code | 8.5K grade-school arithmetic word problems requiring multi-step reasoning. | |
| HARDMath | Harvard | 2024 | paper | code | Graduate applied-math problems requiring asymptotic and approximation methods where leading LLMs score below 45%. | |
| HARP | UCL | 2024 | paper | code | US competition math problems (AMC/AIME/USAMO) with human-written ground-truth solutions. | |
| IMO-Bench | Google DeepMind | 2025 | paper | code | Olympiad-level suite testing final answers, proof-writing, grading, and Lean formal proofs, vetted by IMO medalists. | |
| Lean Workbook | Shanghai AI Lab | 2024 | paper | — | — | About 57K formal-informal Lean4 problem pairs auto-formalized from competition math. |
| LeanDojo | Caltech / NVIDIA | 2023 | paper | code | 98,734 theorems and proofs from Lean mathlib with premise annotations for retrieval-augmented theorem proving. | |
| MATH | UC Berkeley | 2021 | paper | code | 12,500 competition mathematics problems (AMC/AIME level) with step-by-step solutions. | |
| MATH-500 | OpenAI | 2023 | paper | code | 500-problem held-out subset of MATH widely used for LLM evaluation. | |
| MATH-Vision (MATH-V) | CUHK MMLab / Shanghai AI Lab | 2024 | paper | code | 3,040 real competition problems with visual context, spanning 16 mathematical disciplines and five difficulty levels. | |
| MathArena | ETH Zurich / INSAIT | 2025 | paper | code | Evaluates models only on competitions held after their release date to rule out contamination, covering 162 problems from seven competitions plus human-graded IMO proof writing. | |
| MathBench | Shanghai AI Laboratory | 2024 | paper | code | Hierarchical bilingual benchmark spanning arithmetic to college-level theory and application. | |
| MATHCHECK | XJTLU / HKUST | 2024 | paper | code | Checklist benchmark testing task generalization and reasoning robustness beyond end-to-end answer accuracy. | |
| MathOdyssey | NetMind.AI | 2024 | paper | code | 387 expert-crafted high-school, university, and olympiad-level math problems. | |
| MathQA | UW / Allen AI | 2019 | paper | code | 37K multiple-choice math word problems annotated with executable operation programs. | |
| MathVerse | CUHK MMLab / Shanghai AI Lab | 2024 | paper | code | 2,612 visual math problems in six diagram/text variants probing whether MLLMs truly interpret figures. | |
| MathVista | UCLA / University of Washington / Microsoft | 2023 | paper | code | Mathematical reasoning benchmark combining visual contexts (figures, charts, diagrams) with problems. | |
| miniF2F | OpenAI | 2021 | paper | code | 488 olympiad-level (AMC/AIME/IMO) problems formalized across multiple proof assistants. | |
| Omni-MATH | Peking University | 2024 | paper | code | 4,428 olympiad-level competition problems across 33+ subdomains and difficulty tiers. | |
| ProofNet | Yale University | 2023 | paper | code | 371 undergraduate theorems for autoformalization and formal proving in Lean. | |
| Putnam-AXIOM | Stanford | 2025 | paper | — | — | Putnam competition problems plus programmatically perturbed variations giving contamination-resistant unseen instances. |
| PutnamBench | UT Austin | 2024 | paper | code | 1,600+ Putnam competition problems formalized in Lean, Isabelle, and Coq. | |
| We-Math | BUPT / Tencent | 2024 | paper | code | 6.5K visual math problems decomposed into 67 hierarchical knowledge concepts to diagnose reasoning versus memorization. |
Physics olympiad, graduate physics, computational physics, and astronomy.
| Benchmark | Org | Year | Paper | Code | Stars | Description |
|---|---|---|---|---|---|---|
| ABench-Physics | Zhejiang University / Ant Group | 2025 | paper | — | — | 500 hard static plus dynamic-variant physics problems testing reasoning and generalization robustness. |
| Astro-QA | ACMIS Lab | 2025 | paper | code | About 2,700 bilingual astronomy questions across six types spanning astrophysics, celestial mechanics, and astrometry. | |
| AstroMLab 1 | AstroMLab Collaboration | 2024 | paper | — | — | 4,425 astronomy multiple-choice questions from Annual Reviews testing LLM astronomical knowledge. |
| AstroVisBench | NSF-Simons CosmicAI Institute | 2025 | paper | code | Evaluates LLMs on end-to-end astronomy computing workflows and scientific result visualization. | |
| AtmosSci-Bench | HKUST | 2025 | paper | code | Atmospheric-science problems spanning dynamics, physics, hydrology, geophysics, and oceanography via templated questions. | |
| ClimaQA | UC San Diego | 2024 | paper | code | Graduate-level climate science QA generated from textbooks with climate scientists in the loop: 566 expert-validated Gold questions plus 3,000 synthetic Silver questions in MCQ, freeform, and cloze form. | |
| CMPhysBench | Chinese Academy of Sciences (IOP) | 2025 | paper | code | 520+ graduate condensed-matter physics problems with a partial-credit metric; top models score below 30%. | |
| CritPt | Argonne National Laboratory / UIUC (50+ physicists, 30+ institutions) | 2025 | paper | code | 71 composite research-project challenges and 190 modular checkpoints authored by 50+ active physicists, auto-graded via numerical, symbolic, and code evaluations. | |
| DiscoverPhysics | Princeton / NYU / Flatiron Institute / Polymathic AI | 2026 | paper | code | 22 simulated worlds with non-standard physics where an agent proposes initial conditions, observes noisy N-body trajectories, and submits a natural-language law plus a Python implementation. | |
| FEABench | Google Research / Harvard | 2025 | paper | code | Finite-element analysis problems solved end-to-end by driving the COMSOL Multiphysics API: 15 manually verified problems (Gold) plus 200 algorithmically parsed tasks (Large). | |
| Gravity-Bench-v1 | University of Toronto | 2025 | paper | code | Agents plan budgeted observations of simulated binary-star systems to discover the underlying gravitational physics. | |
| HiPhO | CUHK / Shanghai AI Lab | 2025 | paper | code | 13 recent physics olympiad exams with step-level grading and human medal-based comparison for (M)LLMs. | |
| LLM-SRBench | Virginia Tech / CMU | 2025 | paper | code | 239 problems across four science domains testing genuine LLM equation discovery over memorized formulas. | |
| NewtonBench | HKUST | 2025 | paper | code | 324 tasks across twelve physics domains testing LLM agents discovering scientific laws via interactive experimentation. | |
| PHYBench | Peking University | 2025 | paper | code | 500 original physics problems from high-school to olympiad level with expression-edit-distance scoring. | |
| PhysGym | KAUST | 2025 | paper | code | 97 simulated interactive physics-discovery problems with four controlled levels of prior knowledge. | |
| PHYSICS | Yale University | 2025 | paper | code | 1,297 expert-annotated university-level physics problems across six core areas with automated evaluation. | |
| PhysReason | Xi'an Jiaotong University | 2025 | paper | code | 1,200 physics problems with step-level automatic scoring for multi-step physics reasoning. | |
| PRL-Bench | Shanghai Jiao Tong University | 2026 | paper | — | — | End-to-end physics research benchmark built from ~100 recent Physical Review Letters papers across five subfields, each turned into a long-horizon task scored by an LLM-as-judge. |
| TPBench | University of Wisconsin-Madison | 2025 | paper | — | — | 57 theoretical physics problems in high-energy theory and cosmology, undergraduate to research level. |
| UGPhysics | HKUST | 2025 | paper | code | 5,520 bilingual undergraduate physics problems across 13 subjects with rule-based judging. |
Molecular property, reaction, retrosynthesis, safety, and chemical knowledge.
| Benchmark | Org | Year | Paper | Code | Stars | Description |
|---|---|---|---|---|---|---|
| ChemBench | LAMALab, University of Jena (Jablonka group) | 2024 | paper | code | 2,700+ curated questions probing LLM chemical knowledge and reasoning against expert chemists. | |
| ChemCoTBench | IDEA Research / PKU / CUHK | 2025 | paper | code | Step-wise chemical reasoning across molecular understanding, editing, optimization, and reaction prediction. | |
| ChemEval | USTC | 2024 | paper | code | Multi-level chemistry benchmark spanning 42 tasks across four progressive difficulty levels for LLMs. | |
| ChemLLMBench | University of Notre Dame et al. | 2023 | paper | code | Eight chemistry tasks including reaction prediction, retrosynthesis, and molecule captioning for LLMs. | |
| ChemSafetyBench | Peking University | 2024 | paper | — | — | Safety benchmark testing LLM handling of hazardous-chemical property, legality, and synthesis queries. |
| MaCBench | LAMALab (Jena) / IIT Delhi | 2024 | paper | code | 1,100+ image-question pairs probing multimodal LLM limitations across chemistry and materials workflows. | |
| MolCap-Arena | Genentech / UIUC | 2024 | paper | code | Battle-based benchmark scoring 20+ LLMs on molecule captions augmenting molecular property prediction. | |
| MoleculeQA | IDEA Research | 2024 | paper | code | 62K QA pairs over 23K molecules evaluating factual accuracy of LLM molecular comprehension. | |
| MolTextQA | UIUC | 2024 | paper | code | 500K QA pairs over 240K PubChem molecules evaluating molecule-structure-to-text understanding and retrieval. | |
| MOOSE-Chem | NTU Singapore / Shanghai AI Lab | 2024 | paper | code | Tests whether LLMs rediscover unseen chemistry hypotheses from 51 Nature/Science-level papers given background information. | |
| ScholarChemQA | KAUST / Notre Dame | 2024 | paper | code | 40K abstract-derived chemistry research questions exposing LLM limits in comprehending scholarly literature. | |
| SMolInstruct / LlaSMol | OSU NLP Group | 2024 | paper | code | Large-scale instruction dataset of 14 small-molecule chemistry tasks used to train and evaluate LlaSMol. | |
| TOMG-Bench | HK PolyU / Shanghai AI Lab | 2024 | paper | code | Evaluates LLMs on text-based open-domain molecule generation across editing, optimization, and customized generation. |
Crystals, materials property prediction, and materials-science knowledge.
| Benchmark | Org | Year | Paper | Code | Stars | Description |
|---|---|---|---|---|---|---|
| ALDbench | Argonne National Laboratory | 2024 | paper | code | Expert open-ended QA benchmark evaluating LLMs on atomic-layer-deposition synthesis for accuracy and specificity. | |
| AtomWorld | USTC / Shanghai AI Laboratory / UNSW | 2025 | paper | code | Evaluates LLM spatial reasoning on crystalline materials via ten atomic-structure editing actions across four modeling categories over CIF files with verifiable metrics. | |
| LLM4Mat-Bench | Princeton (Vertaix) | 2024 | paper | code | Largest benchmark for LLM materials property prediction over 1.9M crystals and 45 properties. | |
| MaScQA | IIT Delhi (M3RG) | 2024 | paper | code | 650 undergraduate-level materials science questions across 14 categories for evaluating LLM knowledge. | |
| MatCha | CUHK-Shenzhen | 2025 | paper | code | 1,500 expert questions across 21 tasks for materials characterization image understanding by multimodal models. | |
| MatSci-NLP | Mila / Universite de Montreal | 2023 | paper | code | Seven materials-science NLP tasks evaluating language models via a unified text-to-schema approach. | |
| MatSciBench | UCLA | 2025 | paper | code | College-level materials-science reasoning benchmark of 1,340 problems spanning six fields, including multimodal questions. | |
| MatText | LamaLab (Jena) / Intel Labs | 2024 | paper | code | Benchmarking framework evaluating language models on crystal property prediction across nine text representations. | |
| MatViX | Duke University | 2024 | paper | code | Benchmarks vision-language models on multimodal information extraction from visually-rich materials-science articles. | |
| MatVQA | Mila / Universite de Montreal | 2025 | paper | — | — | 1,325 questions testing multimodal models on materials imagery like microscopy and diffraction with multi-step reasoning. |
Genomics, proteins, bioinformatics agents, protocols, and research biology.
| Benchmark | Org | Year | Paper | Code | Stars | Description |
|---|---|---|---|---|---|---|
| BioCoder | Yale University (Gerstein Lab) | 2023 | paper | code | Fuzz-tested benchmark evaluating LLMs on generating bioinformatics code with cross-file dependencies and domain knowledge. | |
| BioLLMBench | UCLA | 2023 | paper | — | — | 2,160 runs evaluating GPT-4, Bard, and LLaMA across 24 bioinformatics tasks and metrics. |
| BioMaze | Peking University | 2025 | paper | code | 5.1K real-research pathway problems testing LLM reasoning over biological pathways and perturbations. | |
| BioML-bench | ScienceMachine | 2025 | paper | code | AI agents build end-to-end biomedical ML pipelines across protein, omics, imaging, and drug tasks. | |
| Biomni | Stanford | 2025 | paper | code | General biomedical agent with Biomni-Eval: 433 instances across 10 reasoning tasks. | |
| BioMysteryBench | Anthropic | 2026 | paper | — | — | 99 expert-written bioinformatics tasks over raw datasets, judged on the final biological conclusion, not the path. |
| BioPlanner | FutureHouse / Francis Crick / Oxford | 2023 | paper | code | BioProt dataset automatically evaluating LLMs on generating biology experimental protocols as executable pseudocode. | |
| BioProBench | Peking University | 2025 | paper | code | Tests LLMs on biological-protocol QA, step ordering, error correction, generation, and reasoning across 556K instances. | |
| BixBench | FutureHouse | 2025 | paper | code | Bioinformatics agents tackle 53 real analysis scenarios with ~300 open-answer research questions. | |
| CellVoyager | Stanford | 2025 | paper | code | CellBench: 76 scRNA-seq studies test agents predicting which analyses the authors performed. | |
| GeneBench-Pro | OpenAI | 2026 | paper | — | — | 129 messy, judgment-intensive computational-biology problems across ten genomics domains on realistic datasets. |
| GeneGPT | NCBI | 2023 | paper | code | Teaches LLMs to call NCBI Web APIs; adds GeneHop, evaluated on GeneTuring. | |
| GeneTuring | Columbia University | 2023 | paper | code | 16 genomics tasks, 1,600 questions probing LLM genomic knowledge and reasoning. | |
| Genome-Bench | Princeton / Stanford | 2025 | paper | — | — | 3,000+ genome-engineering multiple-choice questions mined from expert forum discussions for evaluating LLM reasoning. |
| GenoTEX | UIUC | 2024 | paper | code | Expert-curated gene-expression tasks: dataset selection, preprocessing, and gene-trait statistical analysis. | |
| HealthBench | OpenAI | 2025 | paper | — | — | 5,000 multi-turn physician-crafted health conversations scored by rubric criteria for LLM performance and safety. |
| LAB-Bench | FutureHouse | 2024 | paper | code | 2,400+ MCQs across literature, figures, databases, protocols, and DNA/protein sequence tasks. | |
| LABBench2 | FutureHouse (Edison Scientific) | 2026 | paper | code | Successor to LAB-Bench with ~1,900 biology-research tasks in realistic contexts (literature, figures, protocols, databases), giving a sharp difficulty jump over LAB-Bench. | |
| LifeSciBench | OpenAI | 2026 | paper | — | — | 750 expert-authored free-response life-science research tasks across seven biological domains, rubric-graded. |
| ProteinLMBench | Shanghai Jiao Tong University | 2024 | paper | code | 944 verified MCQs assessing LLM comprehension of protein sequences and descriptions. | |
| VirBench | Anthropic (with NCBI, Broad Institute, Pachter Lab) | 2026 | paper | — | — | 120 curated viral-sequence retrieval queries across ~40 pathogens testing whether LLM agents pull correct ground-truth sequence data from NCBI Virus. |
LLM agents that write research code, run data analyses, attempt autonomous discovery, and conduct ML/AI research.
| Benchmark | Org | Year | Paper | Code | Stars | Description |
|---|---|---|---|---|---|---|
| AAAR-1.0 | Penn State et al. | 2024 | paper | code | Expert-annotated tasks assessing LLMs on equation inference, experiment design, and paper-weakness identification. | |
| AstaBench | Allen AI (AI2) | 2025 | paper | code | 2,400+ problems across eleven benchmarks evaluating agents over the full scientific research pipeline. | |
| BLADE | University of Washington | 2024 | paper | code | Twelve datasets and research questions evaluating agents' analytical decisions against expert ground truth. | |
| CORE-Bench | Princeton | 2024 | paper | code | 270 tasks from 90 papers testing agents on computationally reproducing published scientific results. | |
| Curie | University of Michigan | 2025 | paper | code | Framework and 46-question benchmark for rigorous, automated scientific experimentation across four CS domains. | |
| DiscoveryBench | Allen AI (AI2) | 2024 | paper | code | 264 real plus 903 synthetic tasks for data-driven hypothesis discovery across six domains. | |
| DiscoveryWorld | Allen AI (AI2) | 2024 | paper | code | Simulated environment with 120 tasks requiring full cycles of hypothesis, experiment, and analysis. | |
| DSBench | UT Dallas / Tencent AI Lab | 2024 | paper | code | 540 realistic data-analysis and data-modeling tasks sourced from ModelOff and Kaggle competitions. | |
| EXP-Bench | University of Michigan | 2025 | paper | — | — | Benchmarks agents on designing, implementing, and analyzing end-to-end AI research experiments from publications. |
| HeurekaBench | EPFL (MLBio Lab) | 2026 | paper | code | Builds benchmarks of open-ended research questions grounded in real studies and their code to evaluate end-to-end AI co-scientist agents (instantiated as sc-HeurekaBench in single-cell biology). | |
| InnovatorBench | GAIR-NLP (SJTU) | 2025 | paper | code | Evaluates agents on end-to-end innovative LLM-research tasks like data construction, loss design, and reward design. | |
| LiveIdeaBench | Renmin University of China | 2024 | paper | code | Benchmarks scientific idea generation from minimal keyword context across five dimensions: originality, feasibility, fluency, flexibility, and clarity. | |
| LMR-Bench | UT Dallas | 2025 | paper | code | 28 tasks from 23 NLP papers testing whether agents reproduce masked research code, verified by unit tests. | |
| MLAgentBench | Stanford | 2023 | paper | code | 13 ML experimentation tasks where agents read, write, and execute code to improve performance. | |
| MLE-bench | OpenAI | 2024 | paper | code | 75 Kaggle ML-engineering competitions testing agents against human leaderboards. | |
| MLGym | Meta | 2025 | paper | code | Gym framework with 13 open-ended AI research tasks spanning vision, NLP, RL, and game theory. | |
| MLR-Bench | National University of Singapore | 2025 | paper | code | 201 workshop-derived tasks evaluating agents across idea generation, experimentation, and paper writing. | |
| MLRC-Bench | University of Michigan | 2025 | paper | code | Benchmarks agents on proposing and coding novel methods for seven recent ML research competition problems. | |
| Paper2Code | KAIST / DeepAuto.ai | 2025 | paper | code | Multi-agent framework and benchmark generating functional code repositories that reproduce ML research papers. | |
| PaperBench | OpenAI | 2025 | paper | — | — | Agents replicate 20 ICML 2024 papers from scratch, graded on 8,316 rubric criteria. |
| RE-Bench | METR | 2024 | paper | code | Seven open-ended ML research-engineering environments comparing agents against 61 human experts. | |
| ResearchBench | Shanghai AI Lab / NTU | 2025 | paper | — | — | Benchmarks scientific discovery via inspiration retrieval, hypothesis composition, and ranking across twelve disciplines. |
| ResearchCodeBench | Stanford University | 2025 | paper | code | Challenges LLMs to implement novel contributions from recent ML papers by completing TODO code snippets. | |
| RExBench | Boston University | 2025 | paper | code | Twelve tasks where agents autonomously implement novel research extensions to existing papers' codebases. | |
| SciAgentGym | Fudan NLP Group | 2026 | paper | code | Agentic science benchmark pairing an environment of 1,780 domain-specific tools across natural-science disciplines with a tiered suite from elementary tool actions to long-horizon workflows. | |
| SciCode | UIUC / Princeton / Argonne National Lab | 2024 | paper | code | 80 real research coding problems across 16 natural-science subfields, split into 338 subproblems. | |
| ScienceAgentBench | Ohio State (OSU-NLP) | 2024 | paper | code | 102 data-driven discovery tasks from 44 peer-reviewed papers, evaluating agents that write Python programs. | |
| ScienceWorld | University of Arizona / Microsoft Research / Allen AI (AI2) | 2022 | paper | code | Interactive text environment with 30 elementary-science tasks where agents must run the experiment — melt a substance, test conductivity, breed a plant — instead of reciting the answer. | |
| Scientist-Bench | University of Hong Kong | 2025 | paper | code | Evaluates fully autonomous idea-to-paper research systems across CV, NLP, data mining, and IR against expert papers. | |
| SciReplicate-Bench | King's College London | 2025 | paper | code | Benchmarks agents on reproducing executable code for algorithms described in 36 recent research papers. | |
| SUPER | Allen AI (AI2) | 2024 | paper | code | Evaluates agents on setting up and executing tasks from research code repositories. |
Found a missing benchmark or an error? Edit only data/benchmarks.yaml and open a PR — the README, website, and examples regenerate automatically. See CONTRIBUTING.md for the schema and inclusion bar.
Released under the MIT License. Each entry links to its original paper and code; benchmark metadata is drawn from public sources.