{
  "id": 11219660,
  "title": "# Don’t Trust the Agent’s “Done”: Verify the System State",
  "url": "https://urgent.news/2026/10/01/dont-trust-the-agents-done-verify-the-system-state",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-01T15:15:07.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/taiwildlab_79c1fbf3cc5/-dont-trust-the-agents-done-verify-the-system-state-50ob"
  },
  "original_language": "en",
  "account": "The article emphasizes the need to verify the system state when an AI agent claims that a task has been completed. Just because an agent reports \"Done\" does not guarantee that the expected real-world state has been achieved. Verification is crucial, especially as AI systems gain more execution authority.\n\nThe author presents a model called Claim Authority ↓ Evidence ↓ Verdict, which separates claims made by the AI agent from the evidence needed to verify those claims. The evidence should come from a source capable of observing the property being claimed. For example, for a systemd service, the service manager is the authority, while for a database operation, the database itself may be the authority.\n\nThe distinction between the agent's claim and the actual system state is important. The first describes execution events, while the second describes the resulting system state. One cannot automatically assume that the system is in the desired state merely based on the agent's report.\n\nFurthermore, the author highlights a situation where verification systems must preserve the distinction between false claims and unavailable evidence. If verification systems do not handle this distinction, temporary observability failures can be mistaken for false claims, or accidental confirmation can occur when relevant evidence is unavailable.\n\nThe article suggests starting the verification process in \"shadow mode,\" where the verification system observes claims without influencing execution. This allows for comparison between the evidence generated by the verification system and the checks performed manually by humans. After repeated observations, the verification system can be granted more authority, such as enforcing execution conditions based on its verified claims.\n\nThe simplest implementation starts with verifying a single claim, such as \"Service nginx is active,\" using a trusted authority like systemd. The verification result is then compared to the existing manual process, and this process can be repeated to build a stronger foundation for verification.\n\nThe benefits of this pattern are particularly useful for AI agents with tool access, deployment systems, autonomous remediation, data pipelines, workflow automation, infrastructure operations, and multi-agent systems. The more layers between the agent and the final system state, the more critical it becomes to differentiate between the agent's reported success and the actual existence of the desired outcome.",
  "summary": "AI agents are increasingly allowed to do real work. They deploy applications. They restart services. They write to databases. They call APIs. They trigger automation chains. And when they finish, they usually return something like: “Done.” The problem is that a successful tool call is not necessarily proof that the expected real-world state exists . That distinction becomes increasingly important…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}