Urgent.News

What's breaking now, across thousands of outlets.

AI

How I Killed 63 Flaky Tests in One Week With Claude Code: 5 Lessons

TL;DR I had 63 flaky tests in a Node.js monorepo that made every CI run a coin flip. Over one week I pointed Claude Code at them with a reproduction harness instead of a "please fix this" prompt, and it fixed 58 of them for real, quarantined 5, and taught me a lot about what AI coding agents are actually good at. Spoiler: the bottleneck was never the fixing. It was the diagnosis. ๐Ÿš€ The Problemโ€ฆ

Over a span of one week, a reporter managed to eliminate 63 flaky tests in a Node.js monorepo using Claude Code. Flaky tests, also known as flaky tests, are tests that sometimes pass and sometimes fail unpredictably, causing inconsistencies in the Continuous Integration (CI) pipeline. These tests can lead to increased merge latency, erosion of trust among team members, and inflated retry rates.

The reporter had a total of 4,200 tests distributed across 14 packages within the monorepo, with 63 tests being flaky. The monorepo used Vitest for unit tests, Playwright for browser tests, and a few integration tests that interacted with a local Postgres database. Despite having a decent coverage rate, the CI runs often resulted in red results, with only about 30% success in each run.

The costs associated with flaky tests were significant. Every PR needed an average of 1.7 CI runs to land, wasting considerable time. Additionally, team members began to ignore test failures, leading to a delay in detecting real regressions before they caused more damage. The reporter had previously attempted to fix flaky tests manually, but the process was tedious and time-consuming, with a quarter of the work never getting done.

To solve the issue, the reporter decided to leverage Claude Code's capabilities instead of merely asking it to fix the flaky tests. Instead, they built a proper workflow for the AI coding agent to work with. The process began with measuring the flakiness of the tests. The reporter turned off retries and ran the full test suite 20 times, collecting per-test pass/fail data in a JSONL file. This data served as the foundation for diagnosing the flaky tests.

Next, the reporter created a reproduction harness to verify the agent's fixes. The harness, named 'repro.sh', was a small script that repeated the test execution multiple times and only succeeded if the test passed a certain number of consecutive times. The reporter ran this script 20 times overnight, gathering the flake scores for each test and identifying the problematic ones.

Through this process, Claude Code was able to fix 58 out of the 63 flaky tests, thereby addressing the majority of the issue. The remaining 5 tests were quarantined, meaning they were marked as flaky and required further investigation before being addressed. This approach not only saved the reporter time but also improved the overall reliability of the CI pipeline.

The bottleneck, however, was not the fixing of the flaky tests but the diagnosis process itself. By accurately identifying the flaky tests, the reporter was able to effectively utilize Claude Code's capabilities to streamline the process.

Written by urgent.news from Dev.to's reporting โ€” not their text. Machine-written โ€” may contain errors; check the original before relying on it.

Read the original at dev.to โ†’

More in AI

More from Monday 28 September โ†’