Urgent.News

What's breaking now, across thousands of outlets.

AI

AI models are getting better at training other models, Anthropic study finds

AI models are getting better at training other models, Anthropic study finds

Anthropic's latest study has found early evidence that AI models may be on track to achieve artificial general intelligence (AGI), the hypothetical level of intelligence that would surpass human capabilities across most tasks. The research, titled 'Automated Researchers Can Reliably Mitigate Alignment Failures', was conducted by Chen Yueh-Han, an Anthropic researcher in the company's fellows program.

The study involved creating automated alignment researchers (AARs) using Anthropic's Claude Opus 4.8 model to improve a set of alignment benchmarks, which are measures of a model's ability to avoid misaligned behaviors. The AARs were successful in improving performance on all 10 benchmarks without any decrease in overall performance, suggesting a potential step towards recursive self-improvement in AI.

This comes at a crucial time as Anthropic's rival, OpenAI, is also reportedly making significant strides towards AGI, with CEO Sam Altman stating it could happen by the end of this year, according to a report. Mark Chen, OpenAI's chief research officer, estimated that OpenAI is "80 per cent of the way" there. The key findings of Anthropic's study involved building AARs powered by Claude Opus 4.8, which were tasked with searching for training methods, proposing them, and training target AI models using those methods for about 30 minutes on Nvidia H200 GPUs.

The most effective methods were preserved while ineffective ones were discarded, allowing the automated systems to work faster and at a larger scale. The study also compared the performance of AARs in training models to that of human AI researchers, finding that the best AAR method was on par with, and in some cases superior to, human researchers, within six hours.

Additionally, AARs could be a more cost-effective alternative, costing roughly $4 per hour in API inference compared to $150 per hour paid to human researchers. Despite these promising results, the study also highlighted a few limitations. For instance, AARs may not replace human AI researchers anytime soon as their effectiveness depends on the benchmarks reflecting actual alignment goals.

Establishing and maintaining these benchmarks will require significant work, and expanding AI research literature that automated researchers draw from will also need human input.

Written by urgent.news from The Indian Express's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at indianexpress.com →

More in AI

Prompt Injection in Claude Code Opus 5 Auto Mode

  • Claude Code Opus 5 can execute code with 60-80% success rate from simple summarization request
  • Auto Mode, default for Claude Code since mid-August, blocks 0.00% prompt injection attacks
  • Attack chain demonstrates vulnerability, showing need for layered defenses

More from Sunday 30 August →