Urgent.News

What's breaking now, across thousands of outlets.

AI

Anthropic outlines metrics to track AI development at frontier labs: how much AI R&D is done by AI, how well agents are overseen, and how compute is allocated (Anthropic)

AI systems are becoming exponentially more powerful and have begun to automate more of the process of building themselves.

Anthropic has outlined metrics to assess the pace of AI development within frontier labs. As AI systems become more powerful and begin to automate more of the process used to build them, the public needs more information to understand the progress being made. In this article, Anthropic lays out measurement tools to shed light on three critical aspects of AI development: how much AI is involved in R&D, how well agents are supervised, and how compute is allocated.

The company provides a snapshot of these metrics from within Anthropic. However, if there were coordination on pacing the frontier, as proposed by Anthropic CEO Dario Amodei, these numbers would likely shift. To ensure verification, independent third-party evaluators from multiple organizations will be given access to internal processes, systems, and data, similar to what is available to internal risk assessment teams.

These evaluators will verify safety practices, report incidents, and monitor key metrics such as those outlined in this article.

Anthropic reports these measurements to provide the public, third parties, and governments with better visibility into the pace of AI development within frontier labs. The measurements focus on how models are built, enabling a better understanding of the relationship between model inputs (like compute) and outputs (like capabilities). These metrics complement capability evaluations, which measure what models can do, and are published separately through Anthropic's Responsible Scaling Policy (RSP) risk reports.

The measurements presented in this article are focused on the production process of models, helping to correlate model inputs with outputs. These measures complement capability evaluations and are a starting point for monitoring the pace of AI development from outside the labs. The article also highlights the importance of measuring AI-led R&D, as frontier AI labs increasingly use AI to build future AI models, allowing for faster development and safety testing.

However, models that accelerate their own development could make it more challenging for humans to understand or control these systems. Therefore, sharing these metrics is crucial for understanding how close the world is to achieving recursive self-improvement.

Written by urgent.news from Techmeme's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 3 other outlets

Read the original at anthropic.com →

More in AI

More from Thursday 17 September →