Urgent.News

What's breaking now, across thousands of outlets.

AI

Agentic AI is breaking the token meter, and enterprises need a plan for what comes next

Per-token pricing was the best thing to happen to enterprises looking to experiment with artificial intelligence, but it may be the worst thing for AI in production. That’s the quandary at the center of a new Futurum report, “The Off Ramp From Per-Token Pricing,” sponsored by neocloud provider QumulusAI Inc. The report’s key finding is […] The post Agentic AI is breaking the token meter, and…

Agentic AI is breaking the token meter, and enterprises need a plan for what comes next

Per-token pricing has revolutionized the realm of artificial intelligence experimentation for enterprises, but it presents a major challenge when deploying AI in production environments. A recent report by Futurum, titled "The Off Ramp From Per-Token Pricing," sponsored by QumulusAI Inc., highlights the issue that agentic AI can consume ten to one hundred times more tokens per task than simple inference calls.

The report's key takeaway is that agentic AI is significantly more expensive, a fact that many chief information officers and chief financial officers have repeatedly voiced concerns about.

The allure of per-token pricing is clear. Developers can quickly create prototypes using application programming interfaces without worrying about capacity planning or procurement cycles. However, this model fails to distinguish between pilot projects and production systems serving large numbers of employees. Agentic AI, which plans, calls tools, checks its work, retries, hands off, and summarizes, generates tokens at each step, making it a prime candidate for expensive token consumption.

The Futurum forecast predicts that agent and reasoning inference will grow by 219% this year, with total inference spending rising from $120 billion in 2025 to a staggering $885 billion by 2030. This steep increase in spending, coupled with a linear pricing model, can lead to out-of-control budgets that outpace the value created by the AI project.

Companies that have fully integrated agentic AI into their operations have experienced rapid cost increases, often catching the attention of finance and IT leaders, who then question the utility and affordability of the tool.

The risk extends beyond a high bill, as unpredictable costs can cause valuable projects to be abandoned due to difficulty in forecasting and justifying expenses. This governance failure disguised as a pricing issue can significantly hinder AI adoption. Interestingly, the market has shifted away from the assumption that everything must be hosted in the public cloud.

According to a Futurum survey of 824 AI decision-makers, 66% of AI compute consumption is now on reserved and owned infrastructure, compared to 19% for on-demand cloud. Many organizations are comfortable committing to capacity for AI workloads, recognizing that the decision is no longer solely about whether to commit, but rather which workloads justify such investments.

The report also sheds light on Amberd.ai's deployment on QumulusAI bare metal. By partitioning an eight-GPU Nvidia H200 server into four virtual environments, each with two GPUs, and tiering customers based on latency tolerance, Amberd.ai was able to maximize utilization. Two customers pay for the entire server, while subsequent customers are profitable.

However, it is crucial to note that bare metal environments require extensive custom engineering, limiting the operating margin gains for teams without the necessary hardware expertise. This caveat is often overlooked in the discussion of offramps.

In conclusion, the offramp from per-token pricing offers potential solutions for enterprises, but it requires careful consideration of workloads, infrastructure, and strategic model selection. Open-weight models, which are increasingly good enough for a variety of AI tasks, can be deployed on privately controlled infrastructure, providing cost-effective alternatives for many organizations.

Ultimately, the decision between reserved and per-token pricing should be based on a thorough evaluation of the specific use case, infrastructure capabilities, and strategic goals.

Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at siliconangle.com →

More in AI

More from Monday 28 September →