The benchmark score is the number to trust least
Claude Opus 5.5 scores 58 on Artificial Analysis's Intelligence Index, more than double the median of 26 for comparable reasoning models. That is the figure that gets quoted. It also tells you the least about what the model costs to run. The same evaluation run that produced the 58 recorded everything else. Getting through the Intelligence Index took Opus 5.5 some 260 million output tokens. The…
When Claude Opus 5.5 achieved a score of 58 on Artificial Analysis's Intelligence Index, it was not the most telling figure about the model's costs. That benchmark number received the most attention, but it fails to reveal crucial information about the model's operational expenses. In fact, the same evaluation that yielded the 58 score also captured all other relevant data.
Opus 5.5 required 260 million output tokens to complete the Intelligence Index test, whereas the median model needed only 81 million. This discrepancy is evident in the billing structure, where Artificial Analysis calculates an average cost of $5.98 per task to run the index. Input costs $4 per million tokens, while output costs $20 per million tokens - both at the higher end of the range compared to the median figures of $2 and $10 per million tokens, respectively.
However, the composite score can be misleading. A model that achieves similar results with fewer tokens and a slightly lower index score might actually be more cost-effective for a specific workload. The headline fails to disclose this vital information. Additionally, the score omits latency, another critical factor. Artificial Analysis reports a time to first token of 724.53 seconds for Opus 5.5, significantly higher than the median of 3.77 seconds.
This extended thinking phase, which precedes the first answer token, can greatly impact user experience for interactive tools.
Despite the 58 score placing Opus 5.5 among the top performers in capability, with a large 1M-token context window and support for text-plus-image input, the benchmark should not be the sole criterion for evaluating the model. Instead, treat it as a starting point to narrow down the shortlist, and then conduct your own testing using your own prompts, counting output tokens, calculating the cost per task, and measuring time to first useful token.
If you are specifically considering Opus 5.5, be sure to compare the blended rate quoted by Artificial Analysis, which is $2.94 per million tokens with a 7:2:1 cache-hit/input/output mix, across the nine available API providers before making a final decision.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.