Google found a way to test Gemini without seeing the questions
Growing datasets and public benchmarks are making it harder to tell whether a model is being tested on something it The post Google found a way to test Gemini without seeing the questions appeared first on The New Stack .
Google DeepMind unveiled a novel approach to evaluate its proprietary AI model, Gemini, without exposing its underlying architecture to the evaluators. This double-blind testing methodology ensures that the model weights remain hidden from the evaluators, while the test questions are concealed from Google. The pilot study pitted Gemini 2.5 Flash Lite against private benchmarks from MLCommons and the Singapore AI Safety Institute, focusing primarily on the evaluation process rather than the resulting scores.
Benchmark contamination, which can artificially inflate model performance, has been a concern for many AI models, particularly larger ones. The innovative solution leverages Google Cloud Confidential Space, an NVIDIA H100 Confidential GPU, and Intel TDX host memory encryption to create a secure enclave for the evaluation. The enclave safeguards model weights and evaluation prompts, with encrypted connections transferring the data between the evaluator and the provider.
Remote attestation and code controls further fortify the environment, ensuring that neither party can access or manipulate the other's protected assets. While the system minimizes trust requirements, it cannot eliminate it entirely, as hardware and benchmark management also play a role. Nonetheless, this double-blind evaluation method offers a promising avenue for assessing AI models without the risk of benchmark leakage, ultimately providing more reliable and trustworthy results.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.