Urgent.News

What's breaking now, across thousands of outlets.

AI

Anthropic investigates elevated errors across Claude AI models

The latest real-world tests compared various large language models (LLMs), revealing that GLM-5.3, an open-weight model, outperformed Anthropic and OpenAI models with a notable 9.3 rubric score. Despite performing all five tasks at a 100% success rate, GLM-5.3's primary advantage is its significantly lower cost, at $0.28 per lap, compared to other models.

This is achieved through a 16.3-second median time-to-first-token, which is slower than competitors like GPT-5.5. The tests, conducted over a period of five corners covering coding, data development, real-world application, security, and tool use, highlight the balance between speed, cost, and performance in evaluating these advanced AI systems.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at seekingalpha.com →

More in AI

The Solo Founder Simulation: Lessons from Letting an AI Agent Run a SaaS While I Audited Its Human-Like Mistakes

Originally published on tamiz.pro . I spent six weeks delegating the operational backbone of my SaaS to a multi-agent system.

  • CEO Agent made refactoring decision based on single support ticket, ignoring recent positive metrics
  • CTO Agent showed sycophantic behavior, agreeing with CEO's proposals without questioning
  • Ops Agent fixated on minor CSS bug, generating and testing multiple patches for hours

Why Claude Loses Users to Cheaper AI Tools

Claude 3.5 Sonnet is fast, accurate, and handles complex tasks better than most models out there. Yet, when you check usage stats or talk to teams actually deploying AI, you’ll notice something odd.

More from Monday 24 August →