UK Safety Institute Adopts Shared Evaluation Format to Boost Reproducibility

The UK AI Security Institute is now using EvalEval's open infrastructure to publish evaluation results in a standardized, transparent format. This collaboration builds on earlier joint work and aims to address the lack of reproducibility in AI benchmark reporting. The shared schema and platform are designed to help researchers verify and interpret model performance more reliably.
The release covers five benchmarks from AISI's main experiment—HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0—alongside two cyber-focused evaluations using partially overlapping model sets. Results span six frontier models across the Claude and GPT families. The data accompanies AISI's paper examining how inference-time compute and evaluation protocol shape benchmark performance.
The collaboration traces back to a joint workshop at NeurIPS 2025, where AISI feedback helped shape the Every Eval Ever schema. AISI's broader tooling work includes OptStop for evaluation efficiency, HiBayES for statistical rigor, and standardization efforts in transcript analysis and capability elicitation.
Widespread adoption of standardized evaluation reporting could meaningfully improve how AI capabilities are verified and compared across the industry. Researchers and downstream users may gain greater confidence in benchmark claims, while regulators and auditors could more reliably assess model safety claims. However, the approach's impact depends on whether other labs and institutes adopt the shared schema—without broader participation, the format risks becoming an isolated reference point rather than an industry standard.