Urgent.News

What's breaking now, across thousands of outlets.

More in AI

Same Claude. Different Harness. Very Different Result.

Claude didn’t get smarter. We changed everything around it. Somehow, 6.5 points appeared between them. We beat Claude Code with Claude.

  • Claude's performance increased from 78.9% to 85.4% with system changes
  • System architecture, reliability, and cost-efficiency crucial for model potential
  • Cost difference of 49% between leading and submitted Terminal-Bench runs

More from Friday 18 September →