Same Claude. Different Harness. Very Different Result.
Claude didn’t get smarter. We changed everything around it. Somehow, 6.5 points appeared between them. We beat Claude Code with Claude. Which is a slightly ridiculous sentence, but it is also a useful one. We ran Backboard CLI on Terminal-Bench 2.1 using Claude Opus 4.8 through Amazon Bedrock. Claude Code, using the same model, had a published score of 78.9%. Our submitted result was 85.4% ±…
Claude didn't become smarter; rather, the entire system surrounding it changed, resulting in a 6.5-point difference. The study ran Backboard CLI on Terminal-Bench 2.1 using Claude Opus 4.8 through Amazon Bedrock, achieving an 85.4% score with a ±0.8% margin. The same model was used for Claude Code, which scored 78.9%. This experiment demonstrates that the system's architecture around the model is crucial in unlocking its potential.
Performance was only half of the story; the cost of the run was also a significant factor. The submitted Terminal-Bench run cost $280.72, while the leading run at the time cost $552.67, indicating a 49% cost difference. This benchmark emphasizes the importance of evaluating not just the model but also the system's ability to utilize that model effectively, its reliability, context handling, recovery mechanisms, and cost-efficiency.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.