Urgent.News

What's breaking now, across thousands of outlets.

AI

The cost of proving it works

Drafted with AI help, human-reviewed by The Agent Loop. Short version: Your model bill has a shadow bill: every rollout you run before you trust a number, every judge pass over its output, every trace you keep around. It never arrives as its own line, so nobody budgets it. Arize writes the arithmetic out: production eval cost = traffic volume × sampling rate × evaluation surfaces × evaluator cost…

The cost of proving that a model works is often overlooked, as the expenses are not explicitly accounted for in budgets. Arize identifies the production evaluation cost using the formula: traffic volume × sampling rate × evaluation surfaces × evaluator cost + human review + retention. Four multipliers and two human costs make up this equation.

One data and AI leader revealed that their evaluation cost was 10 times the baseline agent workload. The best GPT-4o agent has a 60% average task success, but pass^8 drops below 25%. This demonstrates that reliability requires rollouts, and the paper calculates the loop cost at around $200 for one trial per task ($0.38 for the agent and $0.23 for simulated user per task).

Judge tuning typically costs about $2,000 instead of the usual $2M. This is achieved through 4,480 configurations, multi-fidelity search, and Alpaca-Eval-style evaluation. With one Alpaca-Eval annotation costing approximately $24, the cost ladder is as follows: deterministic checks first, sampled LLM judges second, and humans for escalation on uncertainty based on consequence.

The shadow bill, which represents the cost of evaluation, is separate from the bill for running the model. Production evaluation can be calculated with the formula: traffic volume × sampling rate × evaluation surfaces × evaluator cost + human review + retention (Arize, 2026).

Offline, the cost can be broken down into dataset size × system variants × evaluator runs × cost per evaluation. During production, the shadow bill is computed with traffic volume × sampling rate × evaluation surfaces × evaluator cost + human review + retention. It is crucial to note that the judge is another inference workload, and retention is storage that is paid for monthly.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Thursday 1 October →