{
  "id": 12393181,
  "title": "The Receipt: I Priced Every Token of One AI Agent Task",
  "url": "https://urgent.news/2026/10/06/the-receipt-i-priced-every-token-of-one-ai-agent-task",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-06T14:16:57.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/karmendra_pandey_43ac6983/the-receipt-i-priced-every-token-of-one-ai-agent-task-4ibp"
  },
  "original_language": "en",
  "account": "A detailed cost breakdown of a single AI agent task reveals surprising insights into the true expenses of using AI models. The example, walking through a customer support triage agent's workflow on AWS Bedrock, shows how every component - from system prompts to retries - contributes to the total cost.\n\nThe raw inference cost of $0.051 might seem low, but the more meaningful figure is the fully-loaded cost of $0.23. This includes the hidden expenses of evaluation harnesses, retrieval infrastructure, and human review times for escalated issues. The receipt underscores that the silent budget killers are retries and unnecessary gates, which can rapidly inflate the bill.\n\nKey learnings from the receipt include recognizing how context (knowledge base retrieval) consumes a significant portion of input tokens, often accounting for 43% of the total. This highlights opportunities for optimization through techniques like better chunking and prompt compression. Additionally, the system prompt's negligible cost once cached demonstrates the value of caching mechanisms to avoid re-pricing operations.\n\nThe receipt serves as the fundamental unit for cost governance in AI systems. Monthly aggregate dashboards provide only a high-level view of spending, while the detailed task-level receipt allows for targeted optimization. Three actionable practices emerge from this approach: 1) Emitting a receipt for each task to understand cost drivers, 2) Logging input/output token counts per span to monitor usage, and 3) Budgeting at the task level with circuit breakers to prevent runaway costs. By pricing eval gates and understanding their true cost, teams can make informed decisions about implementing quality checks versus allowing potential failures.\n\nUltimately, reading the receipt reveals that AI agent costs are not a mystery, but rather a collection of predictable and manageable components. By embracing this granular understanding, organizations can better control expenses, improve efficiency, and make strategic decisions about their AI deployments.",
  "summary": "Everyone talks about AI costs in aggregate. Dashboards show monthly spend, per-model totals, team budgets. Nobody shows you a receipt. So here's one. A single AI agent task, priced line by line — every token, every tool call, every retry. This is a worked example from production-style workloads I describe in my reference architecture work (full paper: https://doi.org/10.5281/zenodo.23178788 ).…",
  "key_points": [
    "$0.051 raw inference cost for AI agent task",
    "Fully-loaded cost of $0.23 includes hidden expenses",
    "Context (knowledge base retrieval) consumes 43% of tokens"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}