{
  "id": 1455920,
  "title": "AI’s trillion dollar token reckoning",
  "url": "https://urgent.news/2026/08/17/ais-trillion-dollar-token-reckoning",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-17T08:26:48.000Z",
  "source": {
    "name": "TechRadar",
    "slug": "techradar",
    "url": "https://www.techradar.com/pro/ais-trillion-dollar-token-reckoning"
  },
  "original_language": "en",
  "account": "The enterprise AI strategy has been fixated on reaching the cutting edge before competitors for two years. This has led to a focus on public cloud accounts, OpenAI or Anthropic API keys, and accepting higher costs for speed. However, the trajectory of AI spending is changing. Gartner projects global AI spending to hit $2.52 trillion by 2026, a 44% increase from the previous year, with $1.37 trillion allocated to AI infrastructure. In mid-2025, procurement of AI entered a \"Trough of Disillusionment,\" where success depends on predictable return on investment rather than innovative pilots. Enterprises are now focusing on sustaining, governing, and defending AI in production. The race to the forefront has ended, and the real challenge lies in the associated costs. Token prices have dropped nearly tenfold since 2021, but overall AI spend by organizations has risen. This is due to more capable models enabling greater ambition. Companies like Anthropic, OpenAI, and Mistral are now offering a range of models, from premium reasoners to more affordable workhorses, as customers refuse to pay flagship prices for every task. CIOs are no longer asking which model to use, but where each workload should run and what the associated costs will be. AI inference costs are rising, and businesses are grappling with the economic implications. For instance, a global bank's AI assistant has resolved over 1.5 million customer inquiries in its first year, but the inference economics are challenging. A single decision can involve five to twenty model calls, each with its own context window, leading to significant costs. Companies are responding to this shift, with Decagon reducing their inference cost per voice query by sixfold after restructuring onto an open-source multi-model stack on NVIDIA Blackwell. The next step is the advent of sub-quadratic attention approaches, which can significantly reduce the cost of long-context reasoning. These advancements will enable large banks to run complex risk modeling, fraud detection, and KYC operations more efficiently. The key to success will not be the cheapest token, but the ability to place compute closest to the data, ensuring compliance with jurisdictional requirements, and implementing effective governance. Enterprises that adopt these strategies will shape the next decade of AI.",
  "summary": "We've now entered AI 2.0, where inference economics, data gravity, and control decide outcomes.",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}