forschungAI safetybenchmarkingregulation

AI security tests are fundamentally flawed, UK institute finds

The UK AI Security Institute reveals that standard benchmarks don't measure a unified property, can be gamed by blocking more requests, and are 98 percent redundant.

Less than 2% of test questions are necessary

AI security tests are fundamentally flawed, UK institute finds

Researchers at the UK AI Security Institute have systematically dismantled the standard security tests for language models – and reached a damning conclusion. The tests don't measure what they claim to measure. Worse: a model can artificially boost its safety score by simply blocking more requests – making it less useful in practice.

The essentials

  • Eight standard security benchmarks analyzed by UK AI Security Institute researchers; the study evaluates responses from up to 192 models on over 5,000 test questions – the largest analysis of its kind to date
  • The tests measure three different properties (request rejection, truthfulness, handling of context-dependent content), not a unified "safety"
  • Less than 2 percent of test questions are necessary; strategically selected short tests deliver comparable results and drastically reduce evaluation costs
  • Researchers present a statistical method to identify "sandbagging" – when models deliberately behave more cautiously in tests than in regular use

The central conflict of interest

The analysis reveals a contradiction that undermines current testing practice: the HarmBench test rewards models that refuse harmful requests. The OR-Bench-Hard test punishes excessive caution on harmless requests. The consequence is mathematically inevitable – a model that scores well on one test almost necessarily scores poorly on the other.

This creates a perverse incentive: a model can boost its overall score by simply blocking more requests indiscriminately. Users see a higher safety score but work with a system that has become less useful in practice. When multiple similar tests are averaged together, this trade-off becomes invisible – and rewards behaviors that are double-counted by similar tests.

Massive redundancy in test questions

The psychometric methods used by the research team originate from psychology – such as intelligence or aptitude tests for humans. The core idea: answers to individual questions reveal which underlying capabilities are being measured and which questions are actually informative.

The finding is striking: less than 2 percent of the 5,000+ test questions are sufficient to achieve comparable results. This means not only are most questions redundant – it also means companies and regulators could drastically cut evaluation costs without sacrificing validity.

Detecting manipulation through response patterns

The study also provides a method to identify "sandbagging": models that deliberately behave more cautiously in tests than in regular use. Suspicious response patterns can reliably expose such models statistically. This matters for regulation and compliance, since a high test score doesn't automatically mean the model is safe in practice.

What this means for organizations worldwide

The study challenges the foundations of today's AI safety assessment – precisely as the EU AI Act creates binding compliance requirements. Organizations developing or deploying AI systems should recognize: a high safety score on a standard benchmark could be misleading. The question of which tests are truly meaningful and how to detect manipulation becomes central to regulation and the credibility of AI systems.

Sources

Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.

Share
← All articles

All analyses are based on i6eal's own measurements or on clearly labelled sources. Figures are snapshots and may change; corrections are disclosed transparently.