{
  "id": 12302805,
  "title": "Why \"Agent Said Done\" and \"System Actually Did It\" Are Not the Same Thing",
  "url": "https://urgent.news/2026/10/06/why-agent-said-done-and-system-actually-did-it-are-not-the-same-thing",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-06T04:45:43.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/opsveritas/why-agent-said-done-and-system-actually-did-it-are-not-the-same-thing-3m23"
  },
  "original_language": "en",
  "account": "When an AI agent publishes a support ticket and reports success with an HTTP 200 response, it doesn't guarantee the ticket actually appears in the ticketing system. This discrepancy highlights a gap most monitoring systems overlook. Agents engaging with real-world systems—such as publishing content, transferring funds, or updating CRM records—make two separate bets: whether the agent generates valid output and whether the desired consequence actually occurs in the real system. Most monitoring primarily tracks the first bet, capturing only the agent's response to an API call (e.g., an HTTP 200). However, the API's success doesn't confirm the operation's real-world impact.\n\nIf an agent queues a ticket for later processing or encounters a downstream validation error, it might still return 200, masking the failure. This issue escalates at scale where thousands of executions occur daily. Even a 1% silent failure rate translates to hundreds of invisible broken promises over a month. The agent isn't malicious; the framework isn't broken. The issue lies in the incomplete contract between the agent and the system.\n\nTo address this, action confirmation is crucial. It involves your own code independently verifying the real-world system's state after the agent's attempt. For instance, after publishing a document, the code should read the document back from the publishing system. After filing a ticket, your code should query the ticketing system for the newly created ticket. These independent checks ensure the system actually reflects the intended change, not just the agent's API response.\n\nMonitoring action confirmation requires a distinct approach. While token/latency checks verify the agent's execution, output parsing checks ensure the response is well-formed. Neither confirms the real-world outcome. Action confirmation, however, involves your code reading the target system and verifying the desired state. This method provides a reliable way to confirm whether the agent truly succeeded or merely reported success. Implementing this requires wiring independent checks to your target system's API, ensuring accurate tracking of whether the agent's actions truly materialized.",
  "summary": "Your AI agent just published a support ticket. HTTP 200. The agent said \"ticket created.\" But when you check the ticketing system, nothing is there. This is the gap most monitoring misses. The Contract Your Agent Makes (and Breaks Silently) When you run an LLM agent that takes real-world actions—publishing content, filing tickets, transferring money, updating a CRM record—you're making two…",
  "key_points": [
    "AI agents report success via HTTP 200, but doesn't guarantee real-world system change.",
    "Monitoring systems often overlook the gap between agent's API response and actual system impact."
  ],
  "editors_take": "The discrepancy between an AI agent's reported success and the actual outcome in the real-world system highlights the need for action confirmation to ensure accurate tracking of the agent's true impact.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}