Google Deepmind announces the first double-blind evaluation of a proprietary frontier KI model – with cryptographic security. The procedure aims to solve a core problem in the KI industry: so-called benchmark contamination, where models have already seen test questions during training and therefore cannot be assessed meaningfully.
The essentials
- Pilot project: Google tests a Gemini Flash Lite model against confidential benchmarks with the Singapore AI Safety Institute
- Technology: Confidential Space from Google Cloud's Confidential Computing portfolio secures data cryptographically
- Key advantage: Neither Google sees the test prompts nor do evaluators see the model weights – technically verifiable for the first time
- Goal: New standard for secure model oversight, especially in sensitive areas like cybersecurity
The dilemma of previous evaluations
Until now, highly sensitive external tests forced the industry into an unsatisfying compromise: either evaluators released their test prompts – then the model provider could see the questions in advance and optimize its model accordingly. Or the provider handed over its model weights – and risked its intellectual property. A current example: delayed evaluations for the ARC-AGI benchmark of Anthropic's Claude 3.5, since the KI company enforces a 30-day data retention period for its strongest models.
Previously, the industry relied on zero-logging protocols and contractual safeguards. Many felt this was insufficient.
How the cryptographic solution works
Google uses Confidential Space – a technology from Google Cloud's Confidential Computing portfolio. It allows external test data and the KI model to run in a cryptographically secured "box." The evaluator does not see the Gemini weights, Google does not see the test prompts. Cryptographic verification ensures that both sides protect their sensitive data – and that the model cannot know the test questions in advance.
Google describes this combination of technical and cryptographic safeguards as a significant advance compared to previous contractual solutions.
Particularly relevant for sensitive areas
According to Google Deepmind, the procedure is especially important for highly sensitive evaluations – such as in cybersecurity or testing by government agencies. Independent organizations could rigorously test advanced models without sacrificing data sovereignty or security.
Google hopes the pilot project will set a new standard for model oversight. The company publishes details on methodology and results in a technical report.
What this means for German companies
The procedure could become relevant for German KI developers and regulators – especially in the context of the EU AI Act, which requires independent evaluations of high-risk systems. If double-blind evaluation becomes the standard, German companies could validate their models more reliably in the future without disclosing trade secrets. At the same time, authorities and independent institutes could conduct more rigorous testing without security concerns.
Sources
Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.




