The 'Godfather of AI' backs a new watchdog plan to track OpenAI and Anthropic's AI risks from the inside
Anthropic and OpenAI said they'd welcome evaluators to track AI risks. A group of watchdogs, backed by pioneers like Geoffrey Hinton, has its answer.
Computer science icon Geoffrey Hinton, often referred to as the "Godfather of AI," joined forces with fellow computer science pioneer Stuart Russell in signing a letter addressed to AI leaders Sam Altman and Dario Amodei of OpenAI and Anthropic. The letter outlined minimum conditions for AI evaluation groups, which recently released their proposal for more effective collaboration between the companies and external risk evaluators.
The AI Evaluator Forum, a coalition of watchdog groups, presented the letter on Friday, seeking embedded third-party evaluators to monitor AI risks from within OpenAI and Anthropic. To maintain credibility, these evaluators must possess scientific objectivity, transparency, independence, and robust safeguards against interference from the companies being assessed.
The letter emphasizes that evaluators should enjoy unfiltered communication with the companies' boards and other oversight bodies, release findings publicly, and be shielded from retaliation while gaining access to the same systems, data, tools, and physical spaces as company assessors. The forum behind the letter includes METR, a nonprofit AI evaluation group that previously examined OpenAI's security incident with Hugging Face, and AVERI, an organization led by former OpenAI employee Miles Brundage and other contributors.
Tech leaders such as Microsoft CEO Satya Nadella have also endorsed the idea of embedded evaluators.
Written by urgent.news from Business Insider's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- White hat hackers just breached OpenAI using Anthropic's Claude in less than 72 hours — and it is a case study in just how fast AI is advancing techradar.com
- OpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot theguardian.com
- Protesters target OpenAI and Anthropic in San Francisco over AI safety fears euronews.com
- Anthropic and OpenAI need truly independent safety evaluators, experts say in public letter cnbc.com
- Security researchers used Anthropic's Claude to hack into OpenAI in under 72 hours qz.com
- Los investigadores logran vulnerar OpenAI utilizando modelos de Anthropic expansion.com
- Chinese AI models fetch fraction of OpenAI, Anthropic revenue despite lower costs: report seekingalpha.com