Urgent.News

What's breaking now, across thousands of outlets.

AI

An Anthropic researcher just gave us a peek at self-improving AI

Given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading overall performance.

Anthropic's researchers have unveiled a groundbreaking study titled "Automated Researchers Can Reliably Mitigate Alignment Failures." In this research, Anthropic Fellow Chen Yueh-Han explored the potential of AI systems to self-improve and address alignment issues. By training AI models with other AI models, the system was able to improve performance on a set of 10 misaligned behaviors without compromising overall performance.

This automated method searched existing literature, proposed methods, and trained the models accordingly. Over time, effective techniques were retained, while ineffective ones were discarded, enabling rapid and scalable improvement.

The paper suggests that these results provide early evidence that automated alignment post-training could become practical in the near future. If models can improve their own alignment training, it raises the possibility that they could enhance training practices broadly, potentially rendering human AI researchers obsolete. The study acknowledges this possibility and directly compares the Automated Alignment Researcher (AAR) to human researchers, stating that the AAR outperforms human researchers on average within six hours.

The cost comparison also favors the automated approach, with the AAR costing around $4 per hour in API inference, compared to the $150 per hour spent on human researchers. However, the paper also highlights certain limitations. The automated system's effectiveness depends on the benchmarks accurately reflecting the actual alignment goals. Additionally, significant work is required to establish and maintain these benchmarks, along with the literature that the automated researchers draw from.

Written by urgent.news from TechCrunch's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at techcrunch.com →

More in AI

Transformers: Understanding the Architecture Behind Modern AI

Introduction Transformers are the heart of modern AI models. AI has seen a lot of breakthrough advancements from ChatGPT to AI, and now.

  • Transformers are the backbone of modern AI models
  • Multi-head attention enables effective semantic similarity capture
  • Positional encodings provide token position information

More from Friday 28 August →