UK AISI and EvalEval Share Reproducible AI Benchmark Results
UK Artificial Intelligence Safety Institute, working with the EvalEval platform, has published verified results for five major benchmarks-HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0-across six frontier language models (Claude Opus 4-4.6, GPT-5-5.4) plus two cyber-focused evaluations, alongside a paper on how inference compute influences performance.
GGLOBAITOOLS DESKSHARE
UK Artificial Intelligence Safety Institute, working with the EvalEval platform, has published verified results for five major ben…
Share this post
Short answer: UK Artificial Intelligence Safety Institute, working with the EvalEval platform, has published verified results for five major benchmarks-HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0-across six frontier language models (Claude Opus 4-4.6, GPT-5-5.4) plus two cyber-focused evaluations, alongside a paper on how inference compute influences performance.
UK AISI EvalEval reproducible AI benchmark results
The UK Artificial Intelligence Safety Institute has begun publishing its latest evaluation findings through the EvalEval platform, aiming to make benchmark outcomes easier to verify and reuse. This move follows earlier joint work that started at a workshop held alongside NeurIPS 2025, where feedback from the institute helped shape a common reporting framework called Every Eval Ever. By adopting that schema and the accompanying Evaluation Cards, the institute is turning its research into a reusable resource for the broader AI community.
Reproducible evaluation matters because the rapid deployment of advanced models creates a growing need for trustworthy evidence about their capabilities and limits. Currently, results appear in many different formats and locations, often missing the details required to run the same test again. Recreating an experiment can be costly, and without clear documentation it is hard to know whether differences in scores stem from genuine model variation or from subtle changes in setup. EvalEval seeks to solve this problem by offering a shared structure for recording evaluations and an open platform where the information needed to interpret them is kept alongside the raw numbers.
In this new phase of the partnership, the safety institute is releasing a set of verified outcomes for five major benchmarks. Those benchmarks are HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0. The data cover six frontier language models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4. In addition, the release includes findings from two cyber-focused evaluations-Cyber CTFs and The Last Ones-which involve a partially overlapping group of models. All of the material accompanies the institute’s recent paper titled How Inference Compute Shapes Frontier LLM Evaluation, which investigates how performance shifts when models are given more processing time at inference and when the testing procedure is altered.
Humanity's Last Exam score variability and Terminal-Bench 2.0 transparency
The paper highlights that scores on Humanity's Last Exam are not fixed; they vary depending on both the evaluation protocol and the amount of compute allowed during inference. The authors present curves that show the proportion of tasks solved as a function of token usage, counting only the first successful attempt for each problem. When an oracle supplies correctness feedback after every try, the models continue to solve additional tasks as they are allowed to use more tokens. By publishing the full experimental context-including the exact configurations, the feedback mechanism, and the token-by-token progress-the institute gives other researchers a concrete reference point for judging whether observed differences are due to the model itself or to the way the test was run.
Similar transparency is provided for Terminal-Bench 2.0, where the institute’s results are placed beside other published evaluations of the same models obtained under alternative conditions. This side-by-side view makes it easier to spot how changes in the evaluation environment influence reported performance, and it offers a baseline for future meta-analyses that aim to aggregate findings across many studies.
The collaboration also outlines concrete steps for different kinds of contributors. Model producers are encouraged to submit verified outcomes that follow the Every Eval Ever template. Benchmark creators can log their test suites and run data using the same schema, ensuring that future comparisons are grounded in a common language. Researchers who study evaluation, governance, or policy can browse the collection of Evaluation Cards to explore trends, spot gaps in reporting, or examine the overall health of the evaluation ecosystem.
EvalEval research community and UK AI Security Institute overview
EvalEval describes itself as a research community dedicated to strengthening the science of model assessment. Its flagship projects include the Every Eval Ever schema, which serves as a shared repository for evaluation results, and Evaluation Cards, which blend benchmark metadata, run specifics, and model details into readable entries. Together, these tools aim to clarify why seemingly similar scores can arise from markedly different experimental conditions.
The UK AI Security Institute, a government-backed body housed within the Department for Science, Innovation and Technology, focuses on understanding the risks posed by powerful AI systems. Its work spans capability analysis, impact assessment, mitigation development, and policy advice. By sharing its evaluation data through an open, standardized channel, the institute hopes to advance both its own mission and the broader goal of building AI that is safe, reliable, and well understood.
Readers interested in the technical details can consult the institute’s paper on inference compute and benchmarking, as well as related publications on hierarchical Bayesian methods for evaluation, early stopping strategies, and the Every Eval Ever and Evaluation Cards resources themselves. The initiative represents a practical step toward making AI evaluation more transparent, comparable, and useful for anyone building or relying on these technologies.
Frequently asked questions
Which benchmarks did the UK AI Safety Institute release results for via EvalEval?
The institute released verified outcomes for five major benchmarks: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0. Also included are two cyber-focused evaluations, Cyber CTFs and The Last Ones.
Which frontier language models were evaluated in the released data?
The data cover six frontier language models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4.
What does the institute’s paper "How Inference Compute Shapes Frontier LLM Evaluation" find about Humanity's Last Exam scores?
The paper shows that scores on Humanity's Last Exam are not fixed; they vary with evaluation protocol and the amount of compute allowed during inference, presenting curves of solved tasks versus token usage and noting that with oracle feedback models solve more tasks as token use increases.
What are Evaluation Cards and how do they support reproducible AI evaluation?
Evaluation Cards blend benchmark metadata, run specifics, and model details into readable entries, part of the EvalEval platform’s Every Eval Ever schema, aiming to clarify why similar scores can arise from different experimental conditions and to provide a shared structure for recording evaluations.
UK AI Security Institute and the EvalEval Coalition have published reproducible benchmark results, releasing Evaluation Cards for five core benchmarks-HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0-tested on six frontier LLMs and two cyber-focused evaluations, using the Every Eval Ever schema to ensure transparency.
Anthropic’s Opus 5.5 and OpenAI’s GPT-6 Sol and Luna models lower token prices-Opus 5.5 at $4 input and $20 output per million tokens (20 % cheaper than Opus 5) and Sol/Luna at $2/$10 and $0.10/$0.50 per million tokens, roughly half the cost of their predecessors-offering developers reduced AI expenses.
Microsoft announced on September 22, 2026 that it helped dismantle the subscription-based AI-driven fraud platform EvilTokens, which had compromised roughly 12,000 Microsoft accounts across about 10,000 organizations worldwide after appearing on Telegram in February 2026.
NO COMMENTS YET
Comments are open. Have a thought or a question? Share it below.