Urgent.News

What's breaking now, across thousands of outlets.

AI

I made CodeRabbit's reviews a third less noisy with an open-source Claude Code skill

AI code review has a noise problem. On a public benchmark of 50 real pull requests, CodeRabbit raised 300 issues. 77 of them were real bugs on the benchmark's list. Most of the rest were noise: things that look like problems but aren't once you read the code. Developers learn to skim review bots, and once you skim, the real bugs slip past you too. I built pr-proof , three Claude Code skills that…

AI code review tools often generate a lot of "noise" in the form of false positives and irrelevant issues. CodeRabbit's reviews of 50 pull requests escalated this problem, raising 300 issues, 77 of which were genuine bugs. Most of the rest were noise, comments that seemed like problems but weren't after examining the code. Developers tend to skim review bots' comments, which can cause real issues to slip through.

To address this, Tanay Kulkarni created three Claude Code skills for pr-proof, which treats every review comment as a claim to be proven against the code before acting on it. When run over the reviews of those 50 PRs, pr-proof kept 72 of the 77 real bugs (93.5%), reduced the noise issues from 223 to 147 (34%), and increased CodeRabbit's F1 score from 35.2% to 40.4% (+5.2 points, 95% CI +1.9 to +8.3).

The core skill, pr-comment-validation, investigates each review comment by reading the entire file the comment refers to, its callers, interfaces, and siblings. It then traces the execution path to check if the described failure actually occurs. The verdict given is valid, partly valid, wrong, or style preference. Two more skills, pr-validation and pr-review, build on this core functionality.

The pr-proof repo is available at https://github.com/TanayK07/pr-proof for anyone to use and verify.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

llama-server's sleep mode loses or crashes on a request that arrives just before it sleeps

TL;DR : llama-server --sleep-idle-seconds N unloads the model after N idle seconds and is documented to reload it for "any new incoming task".

  • Llama-server sleep mode fails or crashes on requests arriving just before sleep
  • 17,653-token prompt remains in queue without response after server falls asleep
  • One-token completion avoids both crash and hang in 123 out of 123 attempts

More from Saturday 3 October →