Anthropic investigates elevated errors across Claude AI models
The latest real-world tests compared various large language models (LLMs), revealing that GLM-5.3, an open-weight model, outperformed Anthropic and OpenAI models with a notable 9.3 rubric score. Despite performing all five tasks at a 100% success rate, GLM-5.3's primary advantage is its significantly lower cost, at $0.28 per lap, compared to other models.
This is achieved through a 16.3-second median time-to-first-token, which is slower than competitors like GPT-5.5. The tests, conducted over a period of five corners covering coding, data development, real-world application, security, and tool use, highlight the balance between speed, cost, and performance in evaluating these advanced AI systems.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.