{
  "id": 8263989,
  "title": "Same Claude. Different Harness. Very Different Result.",
  "url": "https://urgent.news/2026/09/18/same-claude-different-harness-very-different-result",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-18T15:20:23.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/robimbeault/same-claude-different-harness-very-different-result-1k58"
  },
  "original_language": "en",
  "account": "Claude didn't become smarter; rather, the entire system surrounding it changed, resulting in a 6.5-point difference. The study ran Backboard CLI on Terminal-Bench 2.1 using Claude Opus 4.8 through Amazon Bedrock, achieving an 85.4% score with a ±0.8% margin. The same model was used for Claude Code, which scored 78.9%. This experiment demonstrates that the system's architecture around the model is crucial in unlocking its potential. Performance was only half of the story; the cost of the run was also a significant factor. The submitted Terminal-Bench run cost $280.72, while the leading run at the time cost $552.67, indicating a 49% cost difference. This benchmark emphasizes the importance of evaluating not just the model but also the system's ability to utilize that model effectively, its reliability, context handling, recovery mechanisms, and cost-efficiency.",
  "summary": "Claude didn’t get smarter. We changed everything around it. Somehow, 6.5 points appeared between them. We beat Claude Code with Claude. Which is a slightly ridiculous sentence, but it is also a useful one. We ran Backboard CLI on Terminal-Bench 2.1 using Claude Opus 4.8 through Amazon Bedrock. Claude Code, using the same model, had a published score of 78.9%. Our submitted result was 85.4% ±…",
  "key_points": [
    "Claude's performance increased from 78.9% to 85.4% with system changes",
    "System architecture, reliability, and cost-efficiency crucial for model potential",
    "Cost difference of 49% between leading and submitted Terminal-Bench runs"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}