Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet
When Anthropic launched Claude Fable 5.1 this month, it centered the announcement around one benchmark result: its Terminal-Bench-Science score. In The post Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet appeared first on The New Stack .
When Anthropic introduced Claude Fable 5.1 recently, they emphasized a single benchmark result: its Terminal-Bench-Science score. Fable 5.1 achieved 52.6%, while Fable 5 scored 24.7%, more than doubling the latter model. However, Anthropic's conditions for the benchmark score aren't accessible to most users. When tested by the benchmark's leaderboard, Fable 5 achieved 21.4%, and Fable 5.1 wasn't on the independent leaderboard.
I decided to run the benchmark's tasks as a regular user would, focusing on tasks from various scientific fields. The benchmark offers five categories, each with multiple tests. I selected one test per category that could run in a Python environment to compare the models' performance.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.