{
  "id": 10904003,
  "title": "Coverage theatre: your AI hit 90% coverage and still shipped the bug",
  "url": "https://urgent.news/2026/09/30/coverage-theatre-your-ai-hit-90-coverage-and-still-shipped-the-bug",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-30T09:07:33.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/nikolay_chernev_a67479931/coverage-theatre-your-ai-hit-90-coverage-and-still-shipped-the-bug-5bml"
  },
  "original_language": "en",
  "account": "Your AI generated tests reached 90% coverage, but the story has a tragic ending: the shipped code still contained bugs. The problem is that coverage measures which lines of code ran during testing, not whether the tests could detect if those lines were wrong. An AI can quickly churn out tests that hit 90% coverage, but those tests are often one-dimensional happy path cases with no real validation. Just because a line of code was executed doesn't mean the test actually verified the correct behavior of that code. The key question is not \"did this code run in a test?\" but \"would a test fail if this code were broken in some way?\" Coverage and AI output often get the first question right, but they routinely fail the second critical question. The solution is simple but effective: break the code on purpose. Introduce intentional defects, then run the same test suite to see if any tests fail. If the suite stays green despite the defects, the tests are not actually discriminating against bad code. Mutation testing, which systematically introduces defects, is a cheap way to verify the quality of your test suite. Pairing AI-generated tests with a separate verification suite that can catch defects is essential. Don't rely solely on the coverage bar - check the mutation score too. Coverage tells you what code was hit, mutation tells you what your test suite can actually catch. Remember, high coverage is not evidence of good tests, especially when AI is involved. Treat the green coverage bar as a starting point, not a final answer.",
  "summary": "Your AI assistant just generated tests until the coverage bar turned green. 92%. Ship it? Here's the trap: coverage measures which lines ran , not whether a test would notice when they're wrong . AI made that gap free — you can generate 90% coverage in a minute, all of it happy-path, none of it discriminating. That's coverage theatre: it looks like safety, it measures activity, it proves almost…",
  "key_points": [
    "AI-generated tests reached 90% coverage",
    "Coverage measures only executed lines, not correct behavior",
    "Mutation testing verifies test suite quality"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}