Urgent.News

What's breaking now, across thousands of outlets.

AI

Same token price is not the same bill

Claude Sonnet 5.5 and GPT-6.1 Sol shipped a day apart at the exact same price: $2 per million input tokens, $10 per million output. That makes the price sheet the least interesting part of the comparison. I pointed both at a real open source repo (TinyDB) through a small harness and graded 132 agent runs automatically. The first round ended in a 30/30 tie, so I wrote harder tasks and ran them at…

Claude Sonnet 5.5 and GPT-6.1 Sol were released a day apart and priced identically at $2 per million input tokens and $10 per million output. However, the pricing sheet fails to reveal the most crucial difference between the two models. To illustrate this, the author conducted a comprehensive comparison of 132 agent runs using a real open source repository, TinyDB, through a customized harness.

Although the initial results were a 30/30 tie, the author then increased the difficulty of the tasks and ran them at two different effort levels. Surprisingly, the scores remained tied after this adjustment.

Despite the models performing equally well in terms of correctness, there were significant differences in cost, speed, tool use, and code-review tradeoffs. The underlying reason for this disparity lies in the fact that a single token price does not account for the varying number of tokens a model consumes to complete a task. The tokenization process does not consistently count the same text in the same manner across different models, further complicating the comparison.

Moreover, agents exhibit varying levels of tool calls and turns required to achieve a passing result.

The author emphasizes that choosing a coding model based solely on pricing information from the pricing page is misleading, as it compares the wrong unit. To accurately assess the cost effectiveness of a model, one should measure the cost per graded task on their own repository. The author concludes that while the pricing per token may appear identical, the underlying habits and efficiencies of the models differ significantly, leading to different outcomes in real-world applications.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

How I Build and Ship Ten Roblox Games With a Shared Engine and Local AI

Ten Roblox games, a web app and a short-video pipeline, all built and run by one person from a single PC. The stack Text generation runs on Ollama with Qwen3 14B. Images come from Flux.1 through ComfyUI. All of it runs on my own machine, so nothing is billed per piece and publishing doesn't add a fee.

More from Wednesday 7 October →