[{"data":1,"prerenderedAt":30},["ShallowReactive",2],{"nr-en-google-deepmind-doppelblinde-ki-benchmark-evaluierung":3},{"slug":4,"title":5,"dek":6,"date":7,"time":8,"publishedAt":9,"updated":10,"updatedAt":10,"dateFmt":11,"updatedFmt":10,"kind":12,"tier":13,"author":14,"authorName":15,"topics":16,"tracker":22,"trackerLabel":23,"headlineStat":24,"image":25,"ogImage":26,"imageAlt":5,"csv":10,"minutes":27,"words":28,"html":29},"google-deepmind-doppelblinde-ki-benchmark-evaluierung","Google Deepmind Tests First Double-Blind KI Benchmark Evaluation with Cryptography","Google uses Confidential Computing to keep test questions and model weights secret simultaneously. A pilot project with the Singapore AI Safety Institute aims to set a new standard for manipulation-free KI assessments.","2026-08-28","14:19","2026-08-28T14:19:00+02:00","","August 28, 2026","news","standard","ideal-syka","Ideal Syka",[17,18,19,20,21],"KI benchmarks","model evaluation","cryptography","Google Deepmind","trust security","\u002Fstand-der-ki","KI progress","First double-blind evaluation of a proprietary frontier model","\u002Fnewsroom\u002Fimg\u002Fgoogle-deepmind-doppelblinde-ki-benchmark-evaluierung.webp","\u002Fog-nr\u002Fgoogle-deepmind-doppelblinde-ki-benchmark-evaluierung.en.png",2,467,"\u003Cp>Google Deepmind announces the first double-blind evaluation of a proprietary frontier KI model – with cryptographic security. The procedure aims to solve a core problem in the KI industry: so-called \u003Cstrong>benchmark contamination\u003C\u002Fstrong>, where models have already seen test questions during training and therefore cannot be assessed meaningfully.\u003C\u002Fp>\n\u003Ch2>The essentials\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Pilot project\u003C\u002Fstrong>: Google tests a \u003Cstrong>Gemini Flash Lite\u003C\u002Fstrong> model against confidential benchmarks with the \u003Cstrong>Singapore AI Safety Institute\u003C\u002Fstrong>\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Technology\u003C\u002Fstrong>: \u003Cstrong>Confidential Space\u003C\u002Fstrong> from Google Cloud&#39;s Confidential Computing portfolio secures data cryptographically\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Key advantage\u003C\u002Fstrong>: Neither Google sees the test prompts nor do evaluators see the model weights – technically verifiable for the first time\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Goal\u003C\u002Fstrong>: New standard for secure model oversight, especially in sensitive areas like cybersecurity\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>The dilemma of previous evaluations\u003C\u002Fh2>\n\u003Cp>Until now, highly sensitive external tests forced the industry into an unsatisfying compromise: either evaluators released their test prompts – then the model provider could see the questions in advance and optimize its model accordingly. Or the provider handed over its model weights – and risked its intellectual property. A current example: delayed evaluations for the ARC-AGI benchmark of \u003Cstrong>Anthropic&#39;s Claude 3.5\u003C\u002Fstrong>, since the KI company enforces a 30-day data retention period for its strongest models.\u003C\u002Fp>\n\u003Cp>Previously, the industry relied on \u003Cstrong>zero-logging protocols and contractual safeguards\u003C\u002Fstrong>. Many felt this was insufficient.\u003C\u002Fp>\n\u003Ch2>How the cryptographic solution works\u003C\u002Fh2>\n\u003Cp>Google uses \u003Cstrong>Confidential Space\u003C\u002Fstrong> – a technology from Google Cloud&#39;s Confidential Computing portfolio. It allows external test data and the KI model to run in a cryptographically secured &quot;box.&quot; The evaluator does not see the Gemini weights, Google does not see the test prompts. Cryptographic verification ensures that both sides protect their sensitive data – and that the model cannot know the test questions in advance.\u003C\u002Fp>\n\u003Cp>Google describes this combination of technical and cryptographic safeguards as a \u003Cstrong>significant advance\u003C\u002Fstrong> compared to previous contractual solutions.\u003C\u002Fp>\n\u003Ch2>Particularly relevant for sensitive areas\u003C\u002Fh2>\n\u003Cp>According to Google Deepmind, the procedure is especially important for \u003Cstrong>highly sensitive evaluations\u003C\u002Fstrong> – such as in cybersecurity or testing by government agencies. Independent organizations could rigorously test advanced models without sacrificing data sovereignty or security.\u003C\u002Fp>\n\u003Cp>Google hopes the pilot project will set a new standard for model oversight. The company publishes details on methodology and results in a technical report.\u003C\u002Fp>\n\u003Ch2>What this means for German companies\u003C\u002Fh2>\n\u003Cp>The procedure could become relevant for German KI developers and regulators – especially in the context of the EU AI Act, which requires independent evaluations of high-risk systems. If double-blind evaluation becomes the standard, German companies could validate their models more reliably in the future without disclosing trade secrets. At the same time, authorities and independent institutes could conduct more rigorous testing without security concerns.\u003C\u002Fp>\n\u003Ch2>Sources\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.de\u002Fwarum-ki-benchmarks-ein-vertrauensproblem-haben-und-wie-google-es-loesen-will\u002F\">The Decoder (DE)\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cem>Editorially owned by \u003Ca href=\"\u002Fen\u002Fautor\u002Fideal-syka\">Ideal Syka\u003C\u002Fa>. Sources and method: \u003Ca href=\"\u002Fen\u002Fredaktion\">Newsroom &amp; method\u003C\u002Fa>. Tips and corrections: \u003Ca href=\"mailto:ai@i6eal.de\">ai@i6eal.de\u003C\u002Fa>.\u003C\u002Fem>\u003C\u002Fp>\n",1787919990158]