AI agents use 5x more tokens than humans as cached prompts explode, headed for 10x
Daniel Newman’s number comes from OpenRouter data, where agents passed humans in February and grew 14x by August.
Futurum Group CEO Daniel Newman stated that AI agents currently use five times more tokens than humans, a figure that is expected to grow to ten times and beyond. According to data from Andreessen Horowitz (a16z) and OpenRouter, agents consumed 7.3 trillion tokens, compared to humans' 1.4 trillion as of August. However, the majority of these agent tokens consist of cached prompts, with over 85% of them derived from cached prompts, according to a16z's reference to OpenRouter.
The CEO emphasized that human adoption is often overlooked when evaluating AI's Return on Investment (ROI), but the scale and utilization of AI agents are exponentially larger than human usage.
OpenRouter, a leading AI model gateway and routing platform, reported that agent usage surpassed human usage in February and has since grown 14 times, while human usage has increased by 2.8 times. The platform categorizes API keys into three types: agentic, mixed, and human, based on a "7-signal weighted composite score" that includes metrics such as tool call rate, turn count, and gap timing. The mixed category, which may involve a combination of agent and human behavior, grew 4.7 times over the same period.
The data from OpenRouter focuses on token volume rather than spending, and the trend is consistent across multiple platforms, although there have been some fluctuations. In McKinsey's 2026 State of AI survey, 40% of respondents from large organizations reported scaling AI agents, up from 27% a year earlier. Additionally, cached tokens account for nearly all the relative growth in token usage, as a16z cited OpenRouter.
These tokens are less expensive than processing a prompt from scratch, but they require memory storage, which is becoming increasingly costly due to the demand for high-bandwidth memory (HBM).
Models store this context in the KV cache, and the KV cache is outgrowing GPU HBM capacity. This issue is also observed in a call center consultancy's testing of DeepSeek on rented Nvidia H200s, where 96% of input was re-reading old conversation. The data suggests that token counts may overstate the bill, but the hardware cost for memory remains significant.
Micron, a memory manufacturer, anticipates RAM and storage shortages to worsen in 2027 and 2028, with customers facing higher costs while memory makers prioritize HBM for AI data centers. If Newman's prediction that the 5x will become "10X, 20X, 30X" is accurate, PC buyers may face bidding wars for memory against even more AI agents.
Written by urgent.news from Tom's Hardware's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.