[{"data":1,"prerenderedAt":30},["ShallowReactive",2],{"nr-en-uk-ai-security-institute-sicherheitstests-unreliable":3},{"slug":4,"title":5,"dek":6,"date":7,"time":8,"publishedAt":9,"updated":10,"updatedAt":10,"dateFmt":11,"updatedFmt":10,"kind":12,"tier":13,"author":14,"authorName":15,"topics":16,"tracker":22,"trackerLabel":23,"headlineStat":24,"image":25,"ogImage":26,"imageAlt":5,"csv":10,"minutes":27,"words":28,"html":29},"uk-ai-security-institute-sicherheitstests-unreliable","AI security tests are fundamentally flawed, UK institute finds","The UK AI Security Institute reveals that standard benchmarks don't measure a unified property, can be gamed by blocking more requests, and are 98 percent redundant.","2026-08-22","10:14","2026-08-22T10:14:00+02:00","","August 22, 2026","forschung","standard","ideal-syka","Ideal Syka",[17,18,19,20,21],"AI safety","benchmarking","regulation","AI Act","research","\u002Fstand-der-ki","AI progress","Less than 2% of test questions are necessary","\u002Fnewsroom\u002Fimg\u002Fuk-ai-security-institute-sicherheitstests-unreliable.webp","\u002Fog-nr\u002Fuk-ai-security-institute-sicherheitstests-unreliable.en.png",3,504,"\u003Cp>Researchers at the UK AI Security Institute have systematically dismantled the standard security tests for language models – and reached a damning conclusion. The tests don&#39;t measure what they claim to measure. Worse: a model can artificially boost its safety score by simply blocking more requests – making it less useful in practice.\u003C\u002Fp>\n\u003Ch2>The essentials\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Eight standard security benchmarks\u003C\u002Fstrong> analyzed by UK AI Security Institute researchers; the study evaluates responses from \u003Cstrong>up to 192 models\u003C\u002Fstrong> on over \u003Cstrong>5,000 test questions\u003C\u002Fstrong> – the largest analysis of its kind to date\u003C\u002Fli>\n\u003Cli>The tests measure \u003Cstrong>three different properties\u003C\u002Fstrong> (request rejection, truthfulness, handling of context-dependent content), not a unified &quot;safety&quot;\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Less than 2 percent of test questions\u003C\u002Fstrong> are necessary; strategically selected short tests deliver comparable results and drastically reduce evaluation costs\u003C\u002Fli>\n\u003Cli>Researchers present a statistical method to identify \u003Cstrong>&quot;sandbagging&quot;\u003C\u002Fstrong> – when models deliberately behave more cautiously in tests than in regular use\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>The central conflict of interest\u003C\u002Fh2>\n\u003Cp>The analysis reveals a contradiction that undermines current testing practice: the \u003Cstrong>HarmBench test rewards\u003C\u002Fstrong> models that refuse harmful requests. The \u003Cstrong>OR-Bench-Hard test punishes excessive caution\u003C\u002Fstrong> on harmless requests. The consequence is mathematically inevitable – a model that scores well on one test almost necessarily scores poorly on the other.\u003C\u002Fp>\n\u003Cp>This creates a perverse incentive: a model can boost its overall score by simply blocking more requests indiscriminately. Users see a higher safety score but work with a system that has become less useful in practice. When multiple similar tests are averaged together, this trade-off becomes invisible – and rewards behaviors that are double-counted by similar tests.\u003C\u002Fp>\n\u003Ch2>Massive redundancy in test questions\u003C\u002Fh2>\n\u003Cp>The psychometric methods used by the research team originate from psychology – such as intelligence or aptitude tests for humans. The core idea: answers to individual questions reveal which underlying capabilities are being measured and which questions are actually informative.\u003C\u002Fp>\n\u003Cp>The finding is striking: \u003Cstrong>less than 2 percent of the 5,000+ test questions\u003C\u002Fstrong> are sufficient to achieve comparable results. This means not only are most questions redundant – it also means companies and regulators could drastically cut evaluation costs without sacrificing validity.\u003C\u002Fp>\n\u003Ch2>Detecting manipulation through response patterns\u003C\u002Fh2>\n\u003Cp>The study also provides a method to identify &quot;sandbagging&quot;: models that deliberately behave more cautiously in tests than in regular use. Suspicious response patterns can reliably expose such models statistically. This matters for regulation and compliance, since a high test score doesn&#39;t automatically mean the model is safe in practice.\u003C\u002Fp>\n\u003Ch2>What this means for organizations worldwide\u003C\u002Fh2>\n\u003Cp>The study challenges the foundations of today&#39;s AI safety assessment – precisely as the EU AI Act creates binding compliance requirements. Organizations developing or deploying AI systems should recognize: a high safety score on a standard benchmark could be misleading. The question of which tests are truly meaningful and how to detect manipulation becomes central to regulation and the credibility of AI systems.\u003C\u002Fp>\n\u003Ch2>Sources\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.de\u002Fmethoden-aus-der-psychologie-decken-massive-schwaechen-in-ki-sicherheitstests-auf\u002F\">The Decoder (DE)\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.com\u002Fpsychological-methods-reveal-major-weaknesses-in-ai-security-testing\u002F\">The Decoder\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cem>Editorially owned by \u003Ca href=\"\u002Fen\u002Fautor\u002Fideal-syka\">Ideal Syka\u003C\u002Fa>. Sources and method: \u003Ca href=\"\u002Fen\u002Fredaktion\">Newsroom &amp; method\u003C\u002Fa>. Tips and corrections: \u003Ca href=\"mailto:ai@i6eal.de\">ai@i6eal.de\u003C\u002Fa>.\u003C\u002Fem>\u003C\u002Fp>\n",1787386791986]