{
  "id": 13015994,
  "title": "Claude Sonnet 5.5 cache reads cost half: measure your own savings",
  "url": "https://urgent.news/2026/10/09/claude-sonnet-5-5-cache-reads-cost-half-measure-your-own-savings",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-09T03:13:00.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/davekurian/claude-sonnet-55-cache-reads-cost-half-measure-your-own-savings-1202"
  },
  "original_language": "en",
  "account": "On October 7, 2026, Anthropic reduced the cost of reading from the Claude Sonnet 5.5 prompt cache from $0.20 to $0.10 per million tokens. The reduction halves the price of a cache read, but it does not halve the overall model bill for an application. The impact depends on how many cache-read tokens are generated by the traffic, the cost of cache-write tokens, and the frequency of request misses.\n\nTo determine the savings, compare the same workload before and after the rate change, separating the calculation from output tokens, uncached input, retries, and other price changes. Use the application's existing usage counters to calculate the counterfactual change in cache read costs: read_cost_delta = cache_read_tokens / 1,000,000 × ($0.20 - $0.10). For example, a workload generating 25 million cache-read tokens would save $2.50 at the new rate compared to $5 at the previous rate.\n\nThe most significant impact is on routes with substantial, repeated prefixes that already benefit from cache hits. These include stable system instructions, shared reference material, or tool definitions reused across requests. Do not add caching solely based on the lower read cost; instead, measure whether the specific route is reusing the cache prefix. Group traffic by model ID and operation, such as support classification, code review, or a repository assistant's repeated project context.\n\nTo measure the price change, log cache reads, writes, and ordinary tokens separately, using counters for model ID, route, timestamp, request outcome, and internal request class. Calculate the read line and the rate-change delta from one response using the provided code. Retain raw counters for future recomputation if the rates or billing terms change. Also keep track of cache creation and regular input, as a larger read saving can coexist with higher write or uncached-input spend.\n\nWhen comparing like with like, choose a representative window for one stable operation and compare it with another window containing the same operation and a similar request mix. Record all relevant counters and reconcile the token calculation with the provider's billed usage for the same period. Keep the comparison narrow, focusing on \"cache-read spend changed by this amount on this route,\" rather than attributing overall percentage changes to the new read rate.",
  "summary": "Anthropic cut the Claude Sonnet 5.5 prompt-cache read rate on October 7, 2026, from $0.20 to $0.10 per million tokens. The change halves the price of a cache read, but it does not halve an application’s model bill. Your impact depends on how many cache-read tokens your traffic actually generates, how much new caching writes cost, and how often requests miss. Start with the usage counters your…",
  "key_points": [
    "Anthropic reduced Claude Sonnet 5.5 cache read cost to $0.10 per million tokens on October 7, 2026",
    "Greatest impact on routes with repeated prefixes benefiting from cache hits"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}