Urgent.News

What's breaking now, across thousands of outlets.

AI

The cached-prefix crossover: when the cheaper LLM becomes the expensive one

Most LLM cost comparisons collapse a model to one number: $X per million input tokens . That number is the price of a cache miss . Once prompt caching is on, most of the tokens you send on every request are billed at the cached-read rate instead — and caching does not discount every model equally. When it discounts one model far more than another, the ranking flips. That flip has a precise…

When comparing the costs of different language models (LLMs), most analyses only consider a single metric: the price per million input tokens. However, this oversimplification can lead to incorrect conclusions about which model is truly more expensive. The actual cost of using an LLM depends on whether the model benefits from caching—when the model can retrieve frequently used tokens from memory rather than recomputing them.

The impact of caching varies widely across different models. To determine when one model becomes more expensive than another, we can calculate the crossover point, denoted as f*, where the costs of the two models are equal. This occurs when the hit rate (f) satisfies the equation cost_A(f) = cost_B(f). When the cache discount ratio (in / cached) differs significantly between models, the crossover point may shift, causing the cheaper model to change depending on the cache hit rate.

For instance, using a retrieval workload with 50,000 input tokens and 2,000 output tokens per request, it was found that GPT-4.1 becomes more expensive than GPT-6.1 Sol when the hit rate exceeds 53.3%. Conversely, Grok 4.6 and GPT-6.1 Sol switch in favor of each other when the cache hit rate surpasses 40.0%. This discrepancy arises because each model offers different levels of discount on cached reads.

When choosing an LLM for a project, it's crucial to consider not just the headline price, but also the cache discount ratio and how your specific workload will affect the cache hit rate. Failing to do so may result in unexpected and costly invoice discrepancies.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Saturday 10 October →