Research

Frontier AI is Accelerating. Open Benchmarks Need to Keep Up.

Snorkel’s $30M commitment to Open Benchmarks

October 7, 2026
•
7 min read
•

Today, we’re excited to announce that we are expanding our Open Benchmarks Grants by 10x to a $30M commitment, to support researchers, domain experts, and open-source teams pushing the frontier of robust, open frontier AI evaluation. The expanded program includes:

  • Funding the next wave of open benchmarks: More support for new benchmarks and measurement methods including in safety and alignment, cybersecurity, physical AI, long-running agent work, open-ended outputs, dynamic settings, and human uplift.
  • Open Benchmarks Red Team: A new program to help builders continuously identify and address weaknesses in their benchmarks, including reward-hacking exploits, contamination risks, broken verifiers, distributional and diversity gaps, and more.
  • Snorkel Research Fellowship: Support for independent evaluation research, with access to Snorkel’s research team, domain experts, compute, and engineering support.

Through our initial $3M commitment to Open Benchmarks Grants (OBG) and our broader research collaborations, we’ve partnered with the teams behind Terminal-Bench, TB-Science, ARC-AGI-3, OSWorld 2.0, Agents’ Last Exam, Continual Learning Bench, Senior SWE-Bench, SlopCodeBench, and many others. Alongside these teams, we’ve helped design evaluation methodologies, develop task construction pipelines, and scale quality control. Since launch, OBG-funded benchmarks have appeared on the latest model cards from every major frontier lab1; have helped to index and measure frontier progress and alignment; and have contributed to guiding the frontier of AI development.

Benchmarks have always been guideposts for – and drivers of – AI development: setting a metric and then “hillclimbing” against it is the core of how we evaluate and develop AI systems. But when the pace of benchmark development falls behind that of frontier model development, their value for guiding development, evaluation, and alignment of models does as well. “Benchmaxxing” compounds this problem: when a set of benchmarks becomes too simple, static, and correlated, they become too easy to overfit to – like a student studying directly for the questions they expect to see on a simple test that hasn’t been updated – which reduces their value and distorts model development incentives. To mitigate this, we need more benchmarks – that are more robust, diverse, continuously-updated, and independently developed – to serve as a robust, open layer for evaluating and guiding frontier AI.

Massive investments continue to accelerate the pace of AI model development, and we need equally ambitious investment in robust, open measurement. That means funding more diverse benchmarks, continually testing and improving them2, and supporting a broader, more diverse ecosystem of contributors. We’re incredibly excited to contribute to this mission with Open Benchmarks Grants.

Accelerating the frontier of AI evaluation

Closing the evaluation gap requires progress on multiple fronts.

Expanding benchmark diversity and coverage: Each benchmark probes a small region of models’ jagged capabilities. Mapping the broader space requires diverse benchmarks that cover three core axes of AI capability advancement and corresponding benchmark evolution3:

  • Environment and input complexity: Real operating environments are far more complex than today’s benchmark environments. Benchmarks need to capture specialized knowledge, messy context, multimodal inputs, complex toolsets, and coordination with people and other agents.
  • Autonomy horizon: A defining axis of autonomy is how long an agent can operate before reliability breaks down. We need benchmarks that test sustained progress across extended workflows, including whether agents maintain context, recover from mistakes, and adapt as goals and environments change.
  • Output complexity: As agents produce more complex work, evaluation—both of final outputs and of reward signals during training—must become more complex too. Open-ended deliverables require nuanced judgments of quality, usefulness, and appropriate handling of risk.
Diagram of three axes of AI capability advancement: Autonomy Horizon (how independently can the agent operate), Environment Complexity (how dynamic is the operating environment), and Output Complexity (how sophisticated is the deliverable).Diagram of three axes of AI capability advancement: Autonomy Horizon (how independently can the agent operate), Environment Complexity (how dynamic is the operating environment), and Output Complexity (how sophisticated is the deliverable).

Improving measurement robustness: Benchmark scores must remain trustworthy as models improve. Failure modes like contamination, ambiguous tasks, exploitable verifiers, and insufficiently diverse or dynamic data distributions can distort the connection between scores and the capabilities we intend to measure. Maintaining trustworthy measurement is itself a fundamental research challenge, requiring ongoing expert review, adversarial testing, maintenance, and expansion of benchmarks. Open benchmarks let the field inspect methods, reproduce findings, and strengthen tasks / grading mechanisms. Keeping benchmarks trustworthy requires ongoing red-teaming and maintenance from the community— actively probing for weaknesses and addressing them as the frontier moves.

Advancing the empirical science of AI with benchmarks: As AI systems grow more complex and their capabilities become more emergent than engineered, we believe that AI research must become a more empirical science—with benchmarks as one of the key types of controlled experiments. With new emergent capabilities coming into view every day – e.g. from OpenAI’s 10,000-agent mathematics efforts to the covert coordination METR documented in the Hugging Face incident – the importance of developing new evaluation and benchmarking methodologies grows as well. Progress requires researchers with the freedom and support to develop new questions and measurement methods for studying what models can do and whether they operate within their intended goals and constraints.

Expanding Open Benchmarks Grants

Our $30M commitment expands support for independent benchmark builders— academics, domain experts, and open-source teams.

Funding benchmarks where measurement is hardest. We are extending our open invitation for benchmark and evaluation proposals, especially in areas where the stakes are highest and measurement is hardest. These build on and extend the broad axes of interest we laid out at launch – more realistic environments, longer autonomy horizons, more open-ended outputs – and include:

  • Safety & Alignment: Measuring whether agents pursue their intended goals, respect safety constraints, and respond to human oversight.
  • Cybersecurity: Evaluating capabilities and risks in security workflows, from identifying vulnerabilities to investigating incidents.
  • Physical AI: Testing perception, reasoning, and action in physical environments and realistic simulations.
  • Scientific Agents: Measuring agents driving scientific discovery, from hypothesis generation to running experiments.
  • Long-running work: Measuring whether agents can sustain progress, maintain context, and recover from failures across extended workflows.
  • Open-ended Outputs: Grading work with no single correct answer, from research reports to designs, where quality is grounded in expert judgement.
  • Dynamic Settings: Evaluating how agents respond to new information and changing goals in live, event-driven settings.
  • Human Uplift: Measuring how AI assistance changes human performance, relative to working without it.

Introducing the Open Benchmarks Red Team. With this new funding, we’ll be launching a new team to work with benchmark builders on continuously identifying and fixing weaknesses in their evaluations, including reward hacking, contamination, instruction ambiguity/solvability issues, broken verifiers, insufficient distributional diversity, stale task sets, and more. This team will leverage Snorkel’s core data development and evaluation platform and techniques, combining domain expert review with specialized QC and error analysis systems to address failure cases and test fixes. Findings will go to the team first, and we will work with teams to address concrete findings and strengthen their benchmarks.

Supporting researchers through the Snorkel Research Fellowship. We’ll also be launching a new research fellowship program to support active academics and builders working on frontier evaluation. Fellows will set their own research questions and lead work from design through public release, with close collaboration with a Snorkel research co-advisor, along with funded access to domain experts, compute, and research / engineering support.

The work ahead

Benchmarks are guideposts for— and drivers of— AI progress. Keeping pace with the frontier requires more diverse, robust, and continuously evolving benchmarks, built and maintained by a broader ecosystem of researchers and domain leaders. Through Open Benchmarks grants, we’re expanding our support for the researchers, domain experts, and open-source teams building and maintaining these critical tools.

If you have a capability worth measuring or a better way to measure it, we’d like to work with you:

Share this article