358 pull requests that changed tests: agents rarely weakened them. They bent the code instead.
Coding agents get blamed for "making tests pass" by skipping or deleting them. I wanted numbers before building a guard for that, so I collected public pull requests that modified test code (2026, repositories with 100+ stars) and had them labeled before running any tool on them: 198 merged, approved PRs (84 from coding agents, 114 from humans); 98 more merged PRs from an earlier window, held out…
A coding agent's role in modifying test code has been questioned after analyzing 358 pull requests. While most agents did not weaken tests, 2 out of 56 human efforts did remove tests. The majority of intentional changes to checks were due to changes in the behavior being tested, rather than malicious intent. Three of the three instances where agents were found to have weakened tests were due to minor oversights.
The system is designed to flag changes that touch the checks that evaluate the code, but it misses some common simplifications and suppressions. The tool is still in development and can be run locally using the Repopilot repository.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.