Urgent.News

What's breaking now, across thousands of outlets.

Tech

What one agent run actually costs

Drafted with AI help, human-reviewed by The Agent Loop. Short version: You don't have a model bill. You have a loop bill with a model attached, and the loop is where the number comes from. The one number I could actually check I ran the arithmetic on my own session while writing the cost post : 3,518,203 input tokens against 271,350 output tokens. Thirteen tokens read for every one written. The…

The cost of running a single agent can vary greatly, depending on several factors. One key factor is the number of input tokens versus output tokens. In a study of about 4,300 real coding-agent sessions, the median step involved 119,000 tokens from the cached prefix (system prompt, history, tools) and only 875 newly appended tokens.

The model writes approximately 214 tokens per step, making the cached prefix about 136 times larger than the text it appends and roughly 550 times larger than the tokens the model writes. This ratio is per step, not per run, which consists of dozens of stitched-together steps.

Another significant factor is context accumulation. The study found that cached prefix tokens dominated both raw volume and dollar cost in every phase analyzed. As context accumulates, input grows superlinearly, leading to redundant back-and-forth that inflates costs without proportional progress. The cache arithmetic that decides the bill varies by provider, with cache read priced at 0.1x the normal input price and cache write at 1.25x (5-minute cache) or 2x (1-hour cache).

Running multiple steps can lead to cost differences of up to 30 times, with the worst run being about double the cost of the best run for a given problem. Stabilizing the prefix by keeping system prompt, tool schemas, and static history byte-identical across steps can help reduce costs, as a single edit in the prefix can convert reads into writes.

Pruning stale tool output and duplicated file views from the carried history can also help. It's important to measure cost per success rather than per call, as failures often aren't free. Attribution of spend to specific runs, teams, and features at the trace level can provide clarity on costs.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Tailwind CSS joins Shopify; a day later, Shopify drops React Native

Tailwind CSS, the utility-first framework installed over 110 million times a week, is joining Shopify . Twenty-five hours later Shopify published "Native is now the future of mobile at Shopify"…

  • Tailwind CSS joins Shopify after 25 hours
  • Shopify moves apps from React Native to Swift and Kotlin
  • Tailwind Labs reduces staff by 75% due to AI impact

More from Monday 28 September →