{
  "id": 1606295,
  "title": "AI inference is getting cheaper, but your agents are getting more expensive",
  "url": "https://urgent.news/2026/08/18/ai-inference-is-getting-cheaper-but-your-agents-are-getting-more",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-18T01:37:54.000Z",
  "source": {
    "name": "Computerworld",
    "slug": "computerworld",
    "url": "https://www.computerworld.com/article/4210786/ai-inference-is-getting-cheaper-but-your-agents-are-getting-more-expensive.html"
  },
  "original_language": "en",
  "account": "Large language model (LLM) token costs are dropping, yet AI workloads are becoming pricier. Gartner analysts predict that while token costs will plummet by 95% by 2030, inference costs for AI agents will surge more than fivefold in the next two years. This phenomenon, dubbed the \"inference paradox,\" stems from AI app developers utilizing more and pricier tokens as LLMs grow increasingly complex, experts explained. In essence, \"the pace of innovation is outpacing the cost curve,\" noted Gartner analysts Will Sommer and Sabine Zimmerhansl in their report. Buyers may be misled into believing that AI providers' improvements in token economics will translate into cost savings, but this is not the case.\n\nAI proves to be a game-changer, often outperforming humans in routine tasks. Agents can also swiftly identify patterns across fragmented systems. For instance, customer success agents have reduced response times by 99% thanks to AI assistance. However, as AI evolves, token usage escalates, and token value remains unpredictable. Advanced AI agents that can reason already cost up to 150 times more than basic AI chatbots for a single task. More sophisticated agents must engage in reasoning, questioning, and adapting when faced with challenges, and often run continuously in the background, calling upon other agents.\n\nThese requirements demand significant resources, leading to an exponential increase in token consumption, creating a \"massive inference tax\" before users even receive their results. Training medium-sized agentic models with advanced reasoning capabilities costs 2.5 times more than training basic chatbots, while inference costs for agents are 5 times higher, with agents requiring 5 to 30 times more tokens than chatbots to accomplish equivalent tasks.\n\nWhen enterprises deploy hundreds of agents capable of performing dozens or hundreds of tasks per hour, costs skyrocket, and the computational demands become mind-boggling. Gartner developed a Tokenomics Model to assess the impacts of agentic systems, considering various scenarios, including training and inference designs, technology advancements, hardware specifications, and other cost factors. They evaluated 12 types of AI models with varying capabilities and discovered that basic workflows cost around $0.05 per inference token, summarization and knowledge retrieval cost about $0.10, more complex workflows cost around $0.30, and planning and learning tasks cost roughly $0.40 per token. Consequently, the provider cost per token for planning and learning tasks is 8 to 10 times higher than for basic workflows.\n\nTo mitigate these escalating costs, Gartner recommends that enterprises adopt cost-effective strategies. These include developing complex multimodal systems, implementing usage-based pricing, and tiered plans that adjust based on demand. Enterprises should also consider mandating continuous model refresh cycles and setting minimum standards for AI execution, such as defining success thresholds and risk mitigation and compliance measures. By embedding value-per-outcome into product planning, enterprises can forecast and track AI feature spend against clear outcome metrics, helping to identify inefficient or obsolete workflows. Ultimately, ROI from AI advancements is achievable, but it demands substantial effort across intricate workflows, as generic AI solutions can result in exponentially higher costs.",
  "summary": "The good news is that large language model (LLM) token costs are coming down. The conundrum: The overall cost of AI workloads is going up. Gartner research predicts that , while token costs will fall by 95% by 2030, inference costs for agentic workflows will increase more than fivefold over the next two years. This is because AI app builders are using more, and often more expensive, tokens as…",
  "key_points": [
    "LLM token costs dropping 95% by 2030, but inference costs surging more than fivefold",
    "AI agents' complexity leads to inference paradox, where innovation outpaces cost curve",
    "Gartner advises enterprises to adopt cost-effective strategies to mitigate escalating AI costs"
  ],
  "editors_take": "Buyers of AI agents are likely to face rising costs despite falling token costs as complex AI workloads drive up inference expenses, requiring strategic planning to achieve a return on investment.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}