Urgent.News

What's breaking now, across thousands of outlets.

AI

Uber Burned Its Entire 2026 AI Budget by April. Is Your Turn Coming?

Uber Burned Its Entire 2026 AI Budget by April. Is Your Turn Coming? Tokens are the new compute hours, and most teams are managing them like cloud compute -- from 2019. Here's your tokenomics primer. You deployed an AI coding assistant to your engineering team six months ago. Usage lit up immediately. The developers loved it. Then the April invoice arrived. Someone in finance called. The entire…

Uber has exhausted its entire budget for AI in 2026, prompting concerns about the financial challenges that other companies may face in the future. The issue stems from a lack of effective management systems for AI costs. In April, Uber's AI budget for the year was completely spent on deploying an AI coding assistant, Claude Code, to around 5,000 engineers.

This was not a one-off incident, as Microsoft also canceled internal Claude Code licenses due to unexpectedly high token bills. The root cause of these financial troubles lies in the management of AI token consumption, a concept that has gained prominence alongside the decline in per-token inference costs. Economists have termed this phenomenon the Jevons paradox, which suggests that as efficiency improves, total consumption tends to expand, eventually exceeding the initial savings.

The growing trend in token consumption is not following a typical S-curve pattern; instead, it is escalating in a staircase-like manner. Each new architectural pattern, from chat to retrieval-augmented generation to agents, triggers a significant jump in token usage, which briefly stabilizes before jumping again as the next pattern becomes mainstream.

The average length of prompts has increased by roughly 4 times since early 2024, while the average request by developers has increased by nearly 20 times over the same period. These changes have led to a dramatic surge in enterprise generative AI spend, from $1.7 billion in 2023 to $37 billion in 2025. A significant factor contributing to the escalating AI costs is the use of agents.

Agents consume 10 to 100 times more tokens compared to equivalent chat sessions due to the need to resend the full conversation context on every tool call. This behavior can account for a substantial portion of total costs, with one audit finding that 62% of agentic costs were due to re-sent context alone. The situation has become critical, with reports of a single company incurring a $500 million Claude bill in a single month after deploying access without any usage restrictions.

Projected global token usage is expected to multiply by 24 times by 2030, indicating that companies experiencing financial strain in 2026 are still near the beginning of a steep curve. To avoid similar financial crises, organizations need to implement robust observability systems to monitor and manage AI spend effectively. Traditional financial optimization tools often fail to capture the true costs associated with AI, with 70 to 90% of real AI costs going unnoticed.

To address this, organizations must invest in tools that provide detailed breakdowns of costs by model, workflow, and team, along with metrics such as cache hit rates and cache utilization. By gaining a clear understanding of their AI costs, companies can better manage their token economies and prevent unexpected financial setbacks.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Sunday 20 September →