Claude Sonnet 5.5 cache reads cost half: measure your own savings
Anthropic cut the Claude Sonnet 5.5 prompt-cache read rate on October 7, 2026, from $0.20 to $0.10 per million tokens. The change halves the price of a cache read, but it does not halve an application’s model bill. Your impact depends on how many cache-read tokens your traffic actually generates, how much new caching writes cost, and how often requests miss. Start with the usage counters your…
On October 7, 2026, Anthropic reduced the cost of reading from the Claude Sonnet 5.5 prompt cache from $0.20 to $0.10 per million tokens. The reduction halves the price of a cache read, but it does not halve the overall model bill for an application. The impact depends on how many cache-read tokens are generated by the traffic, the cost of cache-write tokens, and the frequency of request misses.
To determine the savings, compare the same workload before and after the rate change, separating the calculation from output tokens, uncached input, retries, and other price changes. Use the application's existing usage counters to calculate the counterfactual change in cache read costs: read_cost_delta = cache_read_tokens / 1,000,000 × ($0.20 - $0.10). For example, a workload generating 25 million cache-read tokens would save $2.50 at the new rate compared to $5 at the previous rate.
The most significant impact is on routes with substantial, repeated prefixes that already benefit from cache hits. These include stable system instructions, shared reference material, or tool definitions reused across requests. Do not add caching solely based on the lower read cost; instead, measure whether the specific route is reusing the cache prefix. Group traffic by model ID and operation, such as support classification, code review, or a repository assistant's repeated project context.
To measure the price change, log cache reads, writes, and ordinary tokens separately, using counters for model ID, route, timestamp, request outcome, and internal request class. Calculate the read line and the rate-change delta from one response using the provided code. Retain raw counters for future recomputation if the rates or billing terms change. Also keep track of cache creation and regular input, as a larger read saving can coexist with higher write or uncached-input spend.
When comparing like with like, choose a representative window for one stable operation and compare it with another window containing the same operation and a similar request mix. Record all relevant counters and reconcile the token calculation with the provider's billed usage for the same period. Keep the comparison narrow, focusing on "cache-read spend changed by this amount on this route," rather than attributing overall percentage changes to the new read rate.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.