DeepMind tests Gemini inside sealed evaluation system
Google DeepMind has completed what it describes as the world’s first double-blind evaluation of a proprietary frontier-class artificial intelligence model, testing Gemini 2.5 Flash-Lite while preventing both the developer and external evaluators from accessing each other’s confidential material. The pilot represents a significant attempt to address benchmark contamination, one of the most…
DeepMind has successfully completed what it claims is the first double-blind evaluation of a proprietary frontier-class AI model, Gemini 2.5 Flash-Lite. In this pilot, the model's developer, Google, and external evaluators were kept separate, preventing either party from accessing confidential information. The evaluation was conducted under strict, secured conditions to protect both the AI model and the test material.
The process utilized Google Cloud's Confidential Space technology and confidential computing protections, ensuring the model and evaluation questions remained isolated and encrypted during processing. Two separate evaluation exercises were performed, with prompts drawn from the MLCommons AILuminate safety benchmark family and a confidential prompt collection supplied by the Singapore AI Safety Institute.
This approach aims to address the issue of benchmark contamination, where model developers may gain access to evaluation questions, or evaluators may require access to proprietary model weights. By keeping the evaluation material and model confidential, the double-blind architecture seeks to protect the interests of both parties - model developers who wish to keep their intellectual property secure, and evaluators who need access to the model without compromising its confidentiality.
While this method does not guarantee the safety of the AI system, it does offer a way to evaluate advanced AI models under conditions that minimize the risk of benchmark contamination and preserve the integrity of the evaluation process.
Written by urgent.news from Arabian Post's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.