UK AISI and EvalEval Publish Reproducible Benchmark Results
UK AI Security Institute and the EvalEval Coalition have published reproducible benchmark results, releasing Evaluation Cards for five core benchmarks-HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0-tested on six frontier LLMs and two cyber-focused evaluations, using the Every Eval Ever schema to ensure transparency.
GGLOBAITOOLS DESKSHARE
UK AI Security Institute and the EvalEval Coalition have published reproducible benchmark results, releasing Evaluation Cards for …
Share this post
Short answer: UK AI Security Institute and the EvalEval Coalition have published reproducible benchmark results, releasing Evaluation Cards for five core benchmarks-HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0-tested on six frontier LLMs and two cyber-focused evaluations, using the Every Eval Ever schema to ensure transparency.
UK AISI EvalEval reproducible benchmark results
The UK AI Security Institute and the EvalEval Coalition have teamed up to make benchmark outcomes more transparent and repeatable. Their collaboration began with a joint workshop held alongside NeurIPS 2025, where feedback from the institute helped shape the Every Eval Ever schema. Now they are putting that shared infrastructure into practice by publishing detailed evaluation data through Evaluation Cards.
Why reproducible evaluation reporting matters for AI systems
Reproducible evaluation reporting matters because assessments of AI systems are becoming a key source of evidence about performance, yet the information needed to repeat those tests is often missing. Results appear in many formats and venues, frequently without the setup details required to run them again. Recreating an evaluation can be costly in time and compute, which limits scrutiny and slows progress. By providing a common schema and an open platform, EvalEval aims to close this gap, while AISI contributes its own work on efficient stopping rules, hierarchical Bayesian methods, and standardised approaches to transcript analysis and capability elicitation. Together they are identifying shortcomings in current reporting and building tools to address them.
As part of this effort, AISI is releasing verified results, context and configuration information for five core benchmarks featured in its recent paper. Those benchmarks are HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The data cover six frontier language models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. In addition, the release includes outcomes from two cyber-focused evaluations, Cyber CTFs and The Last Ones, which involve a partially overlapping set of models. All of this accompanies AISI’s study titled How Inference Compute Shapes Frontier LLM Evaluation, which investigates how performance shifts with the amount of compute used at inference time and with variations in the evaluation protocol.
How evaluation procedure and token budget affect Humanity's Last Exam scores
The paper shows that scores on Humanity's Last Exam depend on both the evaluation procedure and the token budget available to the model. Curves in the study plot the proportion of tasks solved as a function of token usage, counting only the first successful attempt for each task. When an oracle supplies correctness feedback after every try, models continue to solve additional tasks as they are allowed more tokens. This pattern demonstrates how seemingly small changes in setup can alter reported numbers, underscoring the need for full disclosure of methods and conditions.
By publishing the evaluation details alongside the results, AISI gives other researchers the ability to examine individual studies closely and to compare findings across the wider ecosystem. Where other reports omit crucial context, these cards serve as verified reference points that help interpret performance in light of specific choices such as feedback mechanisms or token limits. As more evaluators adopt the Every Eval Ever schema, open comparisons like these can support broader and more reliable meta-research, revealing trends that might be hidden in fragmented disclosures.
How different groups can contribute to AI benchmarking mission
The announcement also outlines how different groups can contribute to the shared mission. Model developers are encouraged to submit verified outcomes from their own tests. Evaluation developers should log benchmarks and run data using the Every Eval Ever framework. Researchers focused on evaluation, governance or policy can explore the cards by benchmark or model, or use the collection to gauge the overall state of reporting practices.
The EvalEval Coalition describes itself as a research community dedicated
Frequently asked questions
What is the purpose of the collaboration between the UK AI Security Institute and the EvalEval Coalition?
The collaboration aims to make benchmark outcomes more transparent and repeatable by providing a common schema (Every Eval Ever) and an open platform, combining AISI's work on efficient stopping rules, hierarchical Bayesian methods, transcript analysis, and capability elicitation with EvalEval's infrastructure to close gaps in evaluation reporting.
Which benchmarks and models are included in AISI's released evaluation data?
AISI released verified results for five core benchmarks-HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0-covering six frontier language models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4, plus outcomes from two cyber-focused evaluations, Cyber CTFs and The Last Ones, which use a partially overlapping set of models.
How does inference compute affect performance on Humanity's Last Exam according to the study?
The study shows that scores on Humanity's Last Exam depend on both the evaluation procedure and the token budget; plotting the proportion of tasks solved versus token usage (counting only the first successful attempt per task) reveals that giving models more tokens lets them solve additional tasks, especially when an oracle supplies correctness feedback after each try.
How can different groups contribute to the EvalEval initiative?
Model developers can submit verified outcomes from their own tests; evaluation developers should log benchmarks and run data using the Every Eval Ever framework; researchers focused on evaluation, governance, or policy can explore the cards by benchmark or model, or use the collection to assess the overall state of reporting practices.
Anthropic’s Opus 5.5 and OpenAI’s GPT-6 Sol and Luna models lower token prices-Opus 5.5 at $4 input and $20 output per million tokens (20 % cheaper than Opus 5) and Sol/Luna at $2/$10 and $0.10/$0.50 per million tokens, roughly half the cost of their predecessors-offering developers reduced AI expenses.
Microsoft announced on September 22, 2026 that it helped dismantle the subscription-based AI-driven fraud platform EvilTokens, which had compromised roughly 12,000 Microsoft accounts across about 10,000 organizations worldwide after appearing on Telegram in February 2026.
Rabbit has launched a cloud-based AI agent called OS3 that runs on any Windows, macOS, or Linux PC without dedicated hardware, letting users connect up to five devices, choose their own AI models, and access the agent via desktop, Telegram, iMessage, or the legacy R1.
NO COMMENTS YET
Comments are open. Have a thought or a question? Share it below.