{
  "id": 11120837,
  "title": "The cost of proving it works",
  "url": "https://urgent.news/2026/10/01/the-cost-of-proving-it-works",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-01T06:04:32.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/theagentloop/the-cost-of-proving-it-works-1jnm"
  },
  "original_language": "en",
  "account": "The cost of proving that a model works is often overlooked, as the expenses are not explicitly accounted for in budgets. Arize identifies the production evaluation cost using the formula: traffic volume × sampling rate × evaluation surfaces × evaluator cost + human review + retention. Four multipliers and two human costs make up this equation.\n\nOne data and AI leader revealed that their evaluation cost was 10 times the baseline agent workload. The best GPT-4o agent has a 60% average task success, but pass^8 drops below 25%. This demonstrates that reliability requires rollouts, and the paper calculates the loop cost at around $200 for one trial per task ($0.38 for the agent and $0.23 for simulated user per task).\n\nJudge tuning typically costs about $2,000 instead of the usual $2M. This is achieved through 4,480 configurations, multi-fidelity search, and Alpaca-Eval-style evaluation. With one Alpaca-Eval annotation costing approximately $24, the cost ladder is as follows: deterministic checks first, sampled LLM judges second, and humans for escalation on uncertainty based on consequence.\n\nThe shadow bill, which represents the cost of evaluation, is separate from the bill for running the model. Production evaluation can be calculated with the formula: traffic volume × sampling rate × evaluation surfaces × evaluator cost + human review + retention (Arize, 2026).\n\nOffline, the cost can be broken down into dataset size × system variants × evaluator runs × cost per evaluation. During production, the shadow bill is computed with traffic volume × sampling rate × evaluation surfaces × evaluator cost + human review + retention. It is crucial to note that the judge is another inference workload, and retention is storage that is paid for monthly.",
  "summary": "Drafted with AI help, human-reviewed by The Agent Loop. Short version: Your model bill has a shadow bill: every rollout you run before you trust a number, every judge pass over its output, every trace you keep around. It never arrives as its own line, so nobody budgets it. Arize writes the arithmetic out: production eval cost = traffic volume × sampling rate × evaluation surfaces × evaluator cost…",
  "key_points": [],
  "editors_take": "The revelation that model evaluation costs can be orders of magnitude higher than expected changes how companies budget for and manage AI model deployment and maintenance.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}