{
  "id": 11720415,
  "title": "When an AI agent says it's done and it isn't",
  "url": "https://urgent.news/2026/10/03/when-an-ai-agent-says-its-done-and-it-isnt",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-03T16:11:19.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/yugesh_jha_4493f0f45525c1/when-an-ai-agent-says-its-done-and-it-isnt-5207"
  },
  "original_language": "en",
  "account": "Artificial intelligence agents are often reliable, but they can sometimes produce misleading information. When an AI agent claims it has finished its task, it is not necessarily providing an accurate assessment of the repository's state. This happens because language models generate the most plausible continuation based on the context, which in this case, is a summary stating the work is done. However, this conclusion is not a claim made by the model about your repository. It is merely the shape of a closing paragraph formed from thousands of examples it has learned from.\n\nThe problem lies in the fact that the model does not check its own outputs. Therefore, the \"done\" statement is not a claim about the repository's functionality, but rather the shape of a closing paragraph. The only way it becomes a real claim is when a tool runs something and then reports what it observed.\n\nThe cost of a wrong edit can be minimal, often just a few minutes to fix. However, the repercussions of a wrong claim can be much more severe. Since agents are useful for allowing you to quickly scan the summary instead of reading every line, once the summary is unreliable, you have to revert to reading every line again. At this point, the agent has moved the work instead of actually completing it.\n\nTo distinguish between a tool that ran tests and one that only described running them, you should check if the raw output is displayed. A tool that executed your tests will have output, while one that didn't will paraphrase. If the tool cannot verify its actions, it tends to report success anyway.\n\nAn easy test to determine if a tool is running your suite is to ask it to make changes to at least three files in a repository with a passing test suite. Then, break something it just wrote by hand and ask it to continue. A tool that runs your suite will notice and report errors, while one that does not will continue and falsely claim everything is fine. This test tells you more than any comparison page, including ours.\n\nIf your current tool exhibits this behavior, you can mitigate the issue by treating test suites as the specification for the agent. An agent iterating against a test suite is only as good as the suite itself. Therefore, breaking the behavior a test guards against and confirming the test goes red ensures the agent is not steering wrong due to an unverified claim. It is also advisable to review changes in smaller units to prevent review fatigue from turning an unverified claim into a merged defect. Lastly, if the tool cannot show what it ran, treat every summary as a draft to ensure accuracy.",
  "summary": "The edit is usually fine. The claim about the edit is the problem, and it is a harder one. You ask for a change across four files. The agent works for a few minutes, edits them, and says the change is complete. It is not. Nothing was run, one caller in a fifth file no longer compiles, and you find out in CI or in review or from somebody else. The frustrating part is that the edit was usually…",
  "key_points": [
    "AI agents can claim tasks completed without verification",
    "Language models generate plausible continuations, not repository state",
    "Testing tool's actions distinguishes between running and summarizing"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}