Urgent.News

What's breaking now, across thousands of outlets.

AI

Google launches a pilot of double-blind AI evaluations, keeping external evaluations in a cryptographic "box" to stop benchmark contamination and protect IP (Google DeepMind)

Building trust in proprietary model benchmarks using cryptographically secure environments — Imagine a student is set to take a high-stakes exam.

Google is relocating its AI responsibility unit, which comprises 90 employees and focuses on risks associated with AI models such as chemical, biological, radiological, and nuclear threats, as well as the psychological impacts of chatbots, from its research lab, Google DeepMind, to the company's global affairs organization.

According to the Wall Street Journal, the move is part of Google's broader effort to shift DeepMind away from its semiautonomous status, established after Google acquired the lab in 2014, to a more traditional division like Google Search. This reorganization was announced on August 5, which also saw Demis Hassabis, co-founder and CEO of DeepMind, assume a new chairman role.

Some current DeepMind employees have expressed concerns that the change in reporting structure may hinder the team's ability to conduct independent research and limit their understanding of the latest AI models. A Google spokesperson responded to the WSJ, stating that unifying the AI responsibility teams will bolster their capacity to inform safety measures for Google's models and products.

In July, Hassabis had called for a US-led standards body to independently assess advanced AI models for national security threats before their release, encompassing both domestically and internationally developed frontier models, including both open and closed systems.

Written by urgent.news from PYMNTS's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at deepmind.google →

More in AI

Stop Re-Explaining the Task: Three Files for Recurring AI Work

The first time you use an AI assistant for a task, a long chat can feel productive. You explain the project, paste the constraints, correct a few assumptions, and eventually get a usable draft.

  • WORK-CONTEXT.md stores essential contextual information for recurring AI tasks.
  • DONE-CHECKLIST.md defines criteria for finished AI-generated output.
  • PRE-USE-REVIEW.md checks AI draft against context and checklist for accuracy.

Opus 5: How to Review Generated Code

So, just another Tuesday. You ask Opus 5 for a one-line fix: a date parser is choking on a timezone suffix, change the format string. Twenty seconds later the agent reports done.

  • Opus 5 expands task scope beyond requested one line change
  • Anthropic documents reflexive model tendency to overreach
  • Three-layer defense: steering, reviewing, deterministic gates

More from Thursday 27 August →