{
  "id": 10463656,
  "title": "How I Killed 63 Flaky Tests in One Week With Claude Code: 5 Lessons",
  "url": "https://urgent.news/2026/09/28/how-i-killed-63-flaky-tests-in-one-week-with-claude-code-5-lessons",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-28T14:31:58.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/yureki_lab/how-i-killed-63-flaky-tests-in-one-week-with-claude-code-5-lessons-29m0"
  },
  "original_language": "en",
  "account": "Over a span of one week, a reporter managed to eliminate 63 flaky tests in a Node.js monorepo using Claude Code. Flaky tests, also known as flaky tests, are tests that sometimes pass and sometimes fail unpredictably, causing inconsistencies in the Continuous Integration (CI) pipeline. These tests can lead to increased merge latency, erosion of trust among team members, and inflated retry rates.\n\nThe reporter had a total of 4,200 tests distributed across 14 packages within the monorepo, with 63 tests being flaky. The monorepo used Vitest for unit tests, Playwright for browser tests, and a few integration tests that interacted with a local Postgres database. Despite having a decent coverage rate, the CI runs often resulted in red results, with only about 30% success in each run.\n\nThe costs associated with flaky tests were significant. Every PR needed an average of 1.7 CI runs to land, wasting considerable time. Additionally, team members began to ignore test failures, leading to a delay in detecting real regressions before they caused more damage. The reporter had previously attempted to fix flaky tests manually, but the process was tedious and time-consuming, with a quarter of the work never getting done.\n\nTo solve the issue, the reporter decided to leverage Claude Code's capabilities instead of merely asking it to fix the flaky tests. Instead, they built a proper workflow for the AI coding agent to work with. The process began with measuring the flakiness of the tests. The reporter turned off retries and ran the full test suite 20 times, collecting per-test pass/fail data in a JSONL file. This data served as the foundation for diagnosing the flaky tests.\n\nNext, the reporter created a reproduction harness to verify the agent's fixes. The harness, named 'repro.sh', was a small script that repeated the test execution multiple times and only succeeded if the test passed a certain number of consecutive times. The reporter ran this script 20 times overnight, gathering the flake scores for each test and identifying the problematic ones.\n\nThrough this process, Claude Code was able to fix 58 out of the 63 flaky tests, thereby addressing the majority of the issue. The remaining 5 tests were quarantined, meaning they were marked as flaky and required further investigation before being addressed. This approach not only saved the reporter time but also improved the overall reliability of the CI pipeline. The bottleneck, however, was not the fixing of the flaky tests but the diagnosis process itself. By accurately identifying the flaky tests, the reporter was able to effectively utilize Claude Code's capabilities to streamline the process.",
  "summary": "TL;DR I had 63 flaky tests in a Node.js monorepo that made every CI run a coin flip. Over one week I pointed Claude Code at them with a reproduction harness instead of a \"please fix this\" prompt, and it fixed 58 of them for real, quarantined 5, and taught me a lot about what AI coding agents are actually good at. Spoiler: the bottleneck was never the fixing. It was the diagnosis. 🚀 The Problem…",
  "key_points": [
    "Reporter eliminated 63 flaky tests in one week using Claude Code",
    "Diagnosed flaky tests by measuring their pass/fail rates over 20 runs",
    "Claude Code fixed 58 out of 63 flaky tests, improving CI pipeline reliability"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}