What is being published
The joint release brings UK AI Security Institute results into Evaluation Cards, a public system built around EvalEval’s Every Eval Ever reporting schema.
It covers five benchmarks from the main experiment:
The results cover six models:
The release also includes two related cyber evaluations, Cyber CTFs and The Last Ones, using a different and partly overlapping set of models.
The data accompanies AISI’s paper How Inference Compute Shapes Frontier LLM Evaluation. Its subject is narrower than model capability in general: how benchmark performance changes with inference resources and with the evaluation protocol itself.
A score is also a protocol
The clearest example is Humanity's Last Exam. A model’s result changes depending on both the evaluation procedure and the amount of inference compute. AISI calculated the cumulative share of attempts solved within a given token budget, counting the earliest recorded success for each task.
Under one protocol, the model received information from an oracle after every attempt about whether its answer was correct. As the token budget grew, it continued solving additional tasks.
That detail changes how the result should be read. A benchmark score is not always a fixed property of a model; it can also describe a particular interaction between the model, the protocol and the available compute.
The same issue appears in the Terminal-Bench 2.0 material, where AISI compares its results with other published evaluations of the same models run under different settings. The comparison does not erase those differences. It makes them visible.
That is the practical value of the release:
I think this is more useful than another leaderboard update. The release does not primarily add a new ranking; it adds the conditions needed to understand why similar-looking rankings may not be measuring the same thing.
The missing standard is still the standard
EvalEval’s Every Eval Ever project provides a common schema and repository for evaluation results. Evaluation Cards combine benchmark metadata, evaluation-run data and model metadata into records intended to remain interpretable.
Together, the projects address a basic reporting gap: results are published in different formats and often omit details needed to repeat an experiment. Re-running an evaluation can also be too expensive to make independent verification routine.
The more important limitation is what the announcement does not claim. It does not say that every evaluation now follows one protocol, that all benchmark results are directly comparable, or that reproducibility is cheap. It offers a shared structure and verified reference points where the details are available.
AISI’s earlier work fits into the same direction:
The collaboration aims to identify gaps in evaluation descriptions and build infrastructure to close them. Its success will depend less on the format itself than on whether model developers and benchmark creators use it consistently.
Who can use it
The organizations invite different groups to contribute:
EvalEval Coalition is a research community developing evidence-based methods and infrastructure for evaluation. Its stated goals include improving evaluation science, creating shared ways to describe the applicability and usefulness of evaluations, and broadening coverage of consequences relevant to research and policy analysis.
UK AI Security Institute is a research organization within the UK Department for Science, Innovation and Technology. It studies the capabilities and consequences of advanced AI, develops and tests safeguards, and provides evidence for policy.
The immediate result is a better-documented set of benchmark runs. The larger test is whether shared reporting can make evaluation results comparable before the next wave of model scores makes the existing inconsistencies even harder to untangle.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X