A comprehensive examination of artificial intelligence safety testing has exposed widespread deficiencies in the benchmarks used to evaluate new AI systems, raising urgent questions about the reliability of technology companies’ claims regarding their products’ safety and capabilities.
Computer scientists from the UK’s AI Security Institute, collaborating with researchers from Stanford University, the University of California Berkeley, and the University of Oxford, analysed more than 440 evaluation benchmarks currently employed to assess AI model safety and effectiveness. The study found that nearly all examined benchmarks contained critical flaws that could render their results “irrelevant or even misleading”.
Fundamental Testing Gaps Identified
The research, led by Andrew Bean from the Oxford Internet Institute, revealed systematic weaknesses across multiple dimensions of AI evaluation methodology.
“Benchmarks underpin nearly all claims about advances in AI,” Bean stated. “But without shared definitions and sound measurement, it becomes hard to know whether models are genuinely improving or just appearing to”.
The investigation uncovered particularly concerning findings regarding statistical rigour. Only 16 per cent of the benchmarks examined employed uncertainty estimates or statistical tests to demonstrate the reliability of their results. This absence of fundamental statistical validation undermines the ability to determine whether performance differences between AI models represent genuine improvements or merely statistical noise.
Additionally, many benchmarks aimed at measuring abstract AI characteristics – such as “harmlessness” or “safety” – relied on vague or contested definitions, making the resulting evaluations less meaningful. This problem of construct validity means that tests may not actually measure what they claim to assess, a flaw that researchers described as particularly troublesome when benchmarks promise to evaluate universal capabilities.
Regulatory Context and Industry Response
The findings arrive at a critical juncture for AI governance. Neither the United States nor the United Kingdom has implemented comprehensive nationwide AI regulation, leaving these benchmarks as principal tools for testing whether new systems are safe, align with human interests, and can perform their claimed capabilities in domains such as reasoning, mathematics, and coding.
The research emerges amid mounting concern over AI safety failures. Google recently withdrew its Gemma AI model from public-facing platforms after it fabricated serious false allegations about US Senator Marsha Blackburn, claiming she had engaged in a non-consensual sexual relationship with a state trooper during a 1987 campaign – an accusation the model supported with fabricated links to non-existent news stories.
“There has never been such an accusation, there is no such individual, and there are no such news stories,” Senator Blackburn wrote in a letter to Google chief executive Sundar Pichai. “A publicly accessible tool that invents false criminal allegations about a sitting US senator represents a catastrophic failure of oversight and ethical responsibility”.
Google clarified that its Gemma models were designed specifically for AI developers and researchers, not for factual assistance or consumer applications. Nevertheless, the incident highlights broader concerns about testing methodologies as AI development accelerates across the industry.
Separate concerns have emerged regarding AI chatbot safety for vulnerable users. Character.AI, a popular AI chatbot platform, recently announced it would ban users under 18 from having open-ended conversations with its AI systems following lawsuits alleging that interactions with the platform contributed to teenage suicides. The company faced legal action from parents who claimed their children confided suicidal thoughts to chatbots that failed to provide appropriate intervention.
Advancing Evaluation Standards
The study’s authors have called for the development of shared standards and best practices to improve AI evaluation processes. The review, conducted by a team of 29 expert reviewers, provides eight key recommendations and detailed guidance for researchers and practitioners developing AI benchmarks.
Among the critical issues identified, researchers highlighted that benchmark practices are fundamentally shaped by cultural, commercial, and competitive dynamics that often prioritise state-of-the-art performance at the expense of broader societal concerns. This creates misaligned incentives where gaming benchmark results becomes more valuable than genuinely improving model capabilities and safety.
Industry efforts to address evaluation challenges include collaborative initiatives. OpenAI and Anthropic recently conducted a pilot evaluation exercise where each company tested the other’s models using internal alignment-related evaluations, marking a rare collaboration between frontier AI developers. The evaluations explored model propensities related to sycophancy, whistleblowing, self-preservation, and supporting human misuse, though both companies emphasised that results should not be interpreted as directly representative of real-world behaviour.
MLCommons, a global consortium, has developed the AI Safety benchmark v0.5 proof-of-concept focusing on measuring the safety of large language models by assessing responses to prompts across multiple hazard categories. The initiative aims to establish industry standards for AI safety evaluation, drawing on multi-institutional expertise to create comprehensive testing frameworks.
Implications for Enterprise and Policy
For organisations making substantial financial commitments to AI programmes—often reaching eight or nine figures—the reliability of benchmark evaluations carries significant implications. Chief technology officers and data officers rely on public leaderboards and benchmarks to compare model capabilities when making procurement and development decisions.
“If a benchmark claiming to measure ‘safety’ or ‘robustness’ doesn’t actually capture those qualities, an organisation could deploy a model that exposes it to serious financial and reputational risk,” noted researchers analysing the construct validity problem.
The European Union has moved forward with regulatory frameworks including the AI Act, which establishes obligations for providers of general-purpose AI models. The Commission confirmed the final General-Purpose AI Code of Practice in July 2025, providing voluntary guidelines to help model providers comply with regulatory requirements. Signatories benefit from a practical grace period until August 2026, though full enforcement of compliance obligations will commence thereafter.
The UK government has strengthened its AI Security Institute (formerly the AI Safety Institute), with a refined focus on advancing understanding of serious risks posed by AI technology and building a scientific evidence base to inform policymakers. The Institute’s research priorities centre on identifying pressing AI security threats, advancing evaluation methods, and creating technical solutions to ensure AI systems remain safe and under human oversight.
Looking Forward
The study’s findings underscore the need for rigorous, transparent evaluation methodologies as AI capabilities continue advancing. Researchers emphasise that without addressing fundamental flaws in current benchmarking practices, the industry risks building systems whose true capabilities and limitations remain poorly understood.
“The flaws undermine the validity of the resulting claims,” the research team concluded, noting that technology companies routinely use these benchmarks to evaluate their latest AI models. As AI systems become increasingly embedded in critical applications spanning healthcare, finance, and public services, the importance of reliable safety evaluation methods will only intensify.
The research calls for urgent reform including improved statistical rigour, clearer construct definitions, enhanced transparency in evaluation processes, and accountability mechanisms to ensure benchmarks genuinely measure what they claim to assess. Only through such improvements can stakeholders – from regulators to enterprises to the general public—place warranted confidence in claims about AI system safety and performance.