Funded benchmarks
Measures how well AI agents design libraries that other agents can use.
Evaluates agents on scientific workflows derived from researchers’ own work.
Evaluates terminal agents on continuously updated software-engineering tasks.
Evaluates coding agents on senior-level software engineering tasks.
Evaluates terminal agents on containerized tasks across seven domains.
Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.
Evaluates computer-use agents on 108 long-horizon workflows.
28 of 89 tasks fixed; continuous validation added to track benchmark reliability.
Measures improvement across sequential, stateful tasks.
Tracks code quality as agents repeatedly extend their own code under changing specs.
Challenges terminal agents with 89 hard, human-verified tasks in containerized environments.
Assesses models on 30 real-world legal tasks using 1,539 rubric scores and 1,530 pairwise judgments from practicing attorneys.
Tests models’ ability to explore and acquire skills in novel environments through hundreds of hand-crafted interactive games.
In development
Assesses AI systems’ capabilities and limitations in English-language philosophy writing.
Tests whether internal activations can reveal covert coordination among AI agents.
Evaluates AI agents’ workbooks against the standards applied to MBA work.
Compares physician trainees and LLMs on clinical relevance across 2,000+ questions, using sentence-level expert labels.
Tests LLM safety in 30-turn conversations with vulnerable users, guided by an adaptive director model.
Measures economically valuable AI capability per unit of energy.

Call for proposals
We are seeking applications from researchers, labs, and engineers building benchmarks for the next wave of AI capabilities. We're looking for benchmarks that drive the fundamental axes for AI agency (read more on our blog) and welcome independent directions from the research community as well.