i
News
News · 2026-09-22

AISI and EvalEval put benchmark conditions on record

@neuronium_ai @neuronium_ai

AISI and EvalEval are publishing reproducible evaluation data for five benchmarks and six frontier models through EvalEval’s Evaluation Cards platform. The release adds tested results, experiment context and configuration details that are often missing from benchmark reports. That matters because model scores can change with the evaluation protocol and the amount of inference compute, making a headline number difficult to interpret—or reproduce—without the conditions behind it.

Cover: AISI and EvalEval put benchmark conditions on record

What is being published

The joint release brings UK AI Security Institute results into Evaluation Cards, a public system built around EvalEval’s Every Eval Ever reporting schema.

It covers five benchmarks from the main experiment:

HealthBench
FrontierMath
Humanity's Last Exam
SWE-Bench Pro
Terminal-Bench 2.0
5benchmarks
6frontier models
2cyber evaluations

The results cover six models:

Claude Opus 4
Claude Opus 4.5
Claude Opus 4.6
GPT-5
GPT-5.2
GPT-5.4

The release also includes two related cyber evaluations, Cyber CTFs and The Last Ones, using a different and partly overlapping set of models.

The data accompanies AISI’s paper How Inference Compute Shapes Frontier LLM Evaluation. Its subject is narrower than model capability in general: how benchmark performance changes with inference resources and with the evaluation protocol itself.

A score is also a protocol

The clearest example is Humanity's Last Exam. A model’s result changes depending on both the evaluation procedure and the amount of inference compute. AISI calculated the cumulative share of attempts solved within a given token budget, counting the earliest recorded success for each task.

Under one protocol, the model received information from an oracle after every attempt about whether its answer was correct. As the token budget grew, it continued solving additional tasks.

That detail changes how the result should be read. A benchmark score is not always a fixed property of a model; it can also describe a particular interaction between the model, the protocol and the available compute.

The same issue appears in the Terminal-Bench 2.0 material, where AISI compares its results with other published evaluations of the same models run under different settings. The comparison does not erase those differences. It makes them visible.

That is the practical value of the release:

Researchers can inspect individual studies rather than relying only on headline scores.
Developers can compare results against a broader set of runs.
Evaluation researchers can see how configuration choices may have affected reported performance.
Policy and governance researchers can study the state of evaluation reporting across benchmarks and models.

I think this is more useful than another leaderboard update. The release does not primarily add a new ranking; it adds the conditions needed to understand why similar-looking rankings may not be measuring the same thing.

The missing standard is still the standard

EvalEval’s Every Eval Ever project provides a common schema and repository for evaluation results. Evaluation Cards combine benchmark metadata, evaluation-run data and model metadata into records intended to remain interpretable.

Together, the projects address a basic reporting gap: results are published in different formats and often omit details needed to repeat an experiment. Re-running an evaluation can also be too expensive to make independent verification routine.

The more important limitation is what the announcement does not claim. It does not say that every evaluation now follows one protocol, that all benchmark results are directly comparable, or that reproducibility is cheap. It offers a shared structure and verified reference points where the details are available.

AISI’s earlier work fits into the same direction:

OptStop makes evaluations more efficient.
HiBayES improves their statistical rigor.
Standardized approaches are applied to transcript analysis and the detection of model capabilities.

The collaboration aims to identify gaps in evaluation descriptions and build infrastructure to close them. Its success will depend less on the format itself than on whether model developers and benchmark creators use it consistently.

Who can use it

The organizations invite different groups to contribute:

Model developers can publish verified evaluation results.
Evaluation developers can describe benchmarks and run data using Every Eval Ever.
Evaluation, governance and policy researchers can study Evaluation Cards by benchmark or model, or analyze evaluation reporting more broadly.

EvalEval Coalition is a research community developing evidence-based methods and infrastructure for evaluation. Its stated goals include improving evaluation science, creating shared ways to describe the applicability and usefulness of evaluations, and broadening coverage of consequences relevant to research and policy analysis.

UK AI Security Institute is a research organization within the UK Department for Science, Innovation and Technology. It studies the capabilities and consequences of advanced AI, develops and tests safeguards, and provides evidence for policy.

The immediate result is a better-documented set of benchmark runs. The larger test is whether shared reporting can make evaluation results comparable before the next wave of model scores makes the existing inconsistencies even harder to untangle.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X