{
  "id": 4251564,
  "title": "Google DeepMind Seals Gemini Test to Protect AI Benchmarks",
  "url": "https://urgent.news/2026/08/28/google-deepmind-seals-gemini-test-to-protect-ai-benchmarks",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-28T21:05:13.000Z",
  "source": {
    "name": "TechRepublic",
    "slug": "techrepublic",
    "url": "https://www.techrepublic.com/article/news-google-deepmind-gemini-tests-apac-singapore/"
  },
  "original_language": "en",
  "account": "Google's AI research arm, DeepMind, has piloted a double-blind evaluation of its proprietary model, Gemini 2.5 Flash Lite, to ensure AI benchmarks remain honest and accurate. Traditional AI testing methods are vulnerable to contamination, where models or their developers may have prior knowledge of test questions, leading to inflated scores. To address this issue, DeepMind employed cryptographic safeguards to protect both the model and evaluation prompts during the testing process. The project involved several organizations, including the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons. The evaluators used reserved prompts from MLCommons' AILuminate safety benchmark, focusing on various harmful content scenarios. The setup involved Google Cloud's Confidential Computing technology, which placed the model and evaluation data inside a protected environment, ensuring that neither party could access the other's sensitive information. This approach could be particularly useful for sensitive evaluations involving cybersecurity and government bodies. However, the pilot did not reveal the actual performance of Gemini 2.5 Flash Lite, as the technical report only described the evaluation architecture and safety categories. While the experiment demonstrates the potential for more secure AI testing, further steps are required to make the process reproducible, transparent, and scalable across models and benchmarks. Until these improvements are made, IT leaders should carefully evaluate vendor claims and consider factors such as the source of benchmarks, evaluation methods, disclosed findings, and the level of trust required in the model provider.",
  "summary": "Google DeepMind tested Gemini 2.5 Flash Lite behind a cryptographic wall designed to protect confidential AI benchmarks and proprietary model weights. The post Google DeepMind Seals Gemini Test to Protect AI Benchmarks appeared first on TechRepublic .",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}