Urgent.News

What's breaking now, across thousands of outlets.

AI

My AI reviewer named the bug 9 of 10 times in a diff and 0 of 10 in raw source. Here is why that is weaker than it sounds.

Limits, before anything else. 21 attempts on one defect are repeated observations, not 21 independent tests. One seeded bug: an inverted authorization check in material I wrote myself. Run on 21 August 2026. One small local model. The alias local-advisor-v1:latest is local. Today ollama show reports architecture qwen35 , 4.7B parameters, Q4_K_M quantization, 8,192-token context, temperature 0.1.…

The AI reviewer identified 9 out of 10 bugs when given the diff version of the code and 0 out of 10 instances when presented with the raw source code. This result is weaker than it may appear at first glance. The defect in question involved an inverted authorization check in the code, allowing users to delete their accounts even if they were not authorized to do so.

The AI reviewer ran the test on August 21, 2026, using a local model named local-advisor-v1:latest, which runs on Ollama with the Qwen35 architecture, 4.7 billion parameters, Q4_K_M quantization, and an 8,192-token context window. The AI reviewer instructed the tool to look for defects related to authorization and correctness. The harness used for the review requested a temperature of 0.1 and schema-constrained JSON output, with no random seed.

Despite the AI reviewer's hand reading the defect in the diff version, none of the 20 valid runs resulted in a successful diagnosis by the tool itself. The tool requires the defect to be named in the findings, along with a file and line reference, and evidence copied from that line. In all valid runs, the findings list remained empty, as the model placed its diagnosis in the free-text control and evidence fields, which the tool labels as unverified model claims.

One significant factor that may have contributed to the difference in results was the --trust-material-paths flag, which instructs the harness that the material is a diff and extracts file paths from it for scope checks. This flag enabled the model to score 9 of 10 diagnoses in the diff version, while the raw source code yielded no diagnoses in any of the 10 runs.

The size of the input material also played a role, as the smallest and largest inputs both scored 0, while the middle-sized diff input scored 9 of 10. The AI reviewer recommends testing reviewers on a seeded defect first, comparing raw-source and diff inputs in actual review workflows, and reporting both the number of diagnoses among valid outputs and successful diagnoses across all attempts, while counting crashes as failures.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Thursday 1 October →