Urgent.News

What's breaking now, across thousands of outlets.

AI

Don’t follow the herd on AI cost optimization – control compute this way

AI cost optimization starts by reducing unnecessary compute, not chasing lower-cost models.

Don’t follow the herd on AI cost optimization – control compute this way

As AI agents move from experimentation to scale, companies are facing higher AI bills. Uber's 2026 AI budget burned through in just four months, prompting interest in model routing and open-source models to cut costs. While switching models can optimize spending, a narrow focus on token usage alone won't be enough. Organizations need to also adopt disciplined methods to determine when to use LLMs and when other data processing techniques are more efficient.

Running LLMs as an analytics engine can lead to unnecessary token consumption since they constantly rebuild context to answer questions. Instead, LLMs should leverage existing organizational knowledge rather than continuously reinventing the wheel. Many problems are repetitive, so models shouldn't need to start from scratch each time.

The most efficient route to AI cost optimization involves integrating LLMs with a business logic layer that contains pre-built workflows, analytics, and compliance rules. This layer can serve as a connective tissue between cloud data platforms and LLMs, avoiding redundant compute costs. By combining trusted workflows, business logic, and targeted use cases, organizations can maximize the value of AI investments and avoid costly mistakes.

Written by urgent.news from TechRadar's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at techradar.com →

More in AI

More from Thursday 1 October →