Urgent.News

What's breaking now, across thousands of outlets.

AI

Why does Opus 5 feel worse to work with?

Some users report that working with Opus 5 feels worse compared to previous models like Opus 4.7, Opus 4.8, and Fable. Despite Opus 5 being a more capable model and even rivaling Fable in benchmarks, it is perceived as less user-friendly. This dissatisfaction stems from the fact that Opus 5 demands more careful oversight and interaction from users.

Two main factors contributing to this issue are Anthropic's ambition to develop a self-improving AI capable of recursively bootstrapping towards AGI/ASI, and the high pressure to achieve high benchmark scores. While it is acknowledged that some benchmark tasks are flawed, a well-defined benchmark should be solvable without requiring hints or outside information.

Selecting models that perform well on benchmarks tends to favor those making bold assumptions in ambiguous situations, discouraging them from asking for clarification. However, this characteristic is precisely what users desire in a coding agent. In real-life scenarios, there is often ambiguity and multiple possible answers, and users prefer an agent that recognizes its limitations and requests further information when necessary.

Real-life situations do not align with the structure of a benchmark, where there is a single correct answer.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at mun-logadan.github.io →

More in AI

More from Friday 14 August →