Funded benchmarks

Sort: Newest
LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

DARPA, National Science Foundation, Laude Institute, Snorkel AI
Link for LibraryDesignBench
Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

Stanford, Laude Institute, Harbor, Snorkel AI
Link for Terminal-Bench-Science
Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

Harbor, Laude Institute, Snorkel AI
Link for Terminal-Bench 4.0
Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Snorkel AI, Princeton University, UW-Madison
Link for Senior SWE-bench
Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

Harbor, Laude Institute, Snorkel AI
Link for Terminal-Bench 3.0
Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

UC Berkeley RDI, RDI Foundation, Snorkel AI
Link for Agents’ Last Exam
OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

XLANG Lab (University of Hong Kong), Snorkel AI
Link for OSWorld 2.0
Terminal-Bench 2.1

28 of 89 tasks fixed; continuous validation added to track benchmark reliability.

Stanford, Laude Institute, Harbor, Snorkel AI
Link for Terminal-Bench 2.1
Continual Learning Bench

Measures improvement across sequential, stateful tasks.

UC Berkeley SkyLab, UW-Madison, Snorkel AI
Link for Continual Learning Bench
SlopCodeBench

Tracks code quality as agents repeatedly extend their own code under changing specs.

UW-Madison, DARPA, National Science Foundation, Snorkel AI
Link for SlopCodeBench
Terminal-Bench 2.0

Challenges terminal agents with 89 hard, human-verified tasks in containerized environments.

Stanford, Laude Institute, Harbor, Snorkel AI
Link for Terminal-Bench 2.0
JudgmentBench

Assesses models on 30 real-world legal tasks using 1,539 rubric scores and 1,530 pairwise judgments from practicing attorneys.

Stanford Law (LIFT Lab), Snorkel AI
Link for JudgmentBench
ARC-AGI-3

Tests models’ ability to explore and acquire skills in novel environments through hundreds of hand-crafted interactive games.

ARC Prize Foundation, Snorkel AI
Link for ARC-AGI-3

In development

PhilosophyBench

Assesses AI systems’ capabilities and limitations in English-language philosophy writing.

Stanford, Snorkel AI
Link for PhilosophyBench
CollusionBench

Tests whether internal activations can reveal covert coordination among AI agents.

Carnegie Mellon University, Snorkel AI
MBABench

Evaluates AI agents’ workbooks against the standards applied to MBA work.

Columbia Business School, Snorkel AI
Link for MBABench
MedPAIR

Compares physician trainees and LLMs on clinical relevance across 2,000+ questions, using sentence-level expert labels.

MIT, Snorkel AI
STELLA-Bench

Tests LLM safety in 30-turn conversations with vulnerable users, guided by an adaptive director model.

Atella AI, Snorkel AI
IntelligencePerWatt

Measures economically valuable AI capability per unit of energy.

Stanford, Snorkel AI
Image

Call for proposals

We are seeking applications from researchers, labs, and engineers building benchmarks for the next wave of AI capabilities. We're looking for benchmarks that drive the fundamental axes for AI agency (read more on our blog) and welcome independent directions from the research community as well.

Apply for a grant