{
  "id": 11241290,
  "title": "My AI reviewer named the bug 9 of 10 times in a diff and 0 of 10 in raw source. Here is why that is weaker than it sounds.",
  "url": "https://urgent.news/2026/10/01/my-ai-reviewer-named-the-bug-9-of-10-times-in-a-diff-and-0-of-10-in",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-01T17:14:29.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/ahmadammar/my-ai-reviewer-named-the-bug-9-of-10-times-in-a-diff-and-0-of-10-in-raw-source-here-is-why-that-is-57mp"
  },
  "original_language": "en",
  "account": "The AI reviewer identified 9 out of 10 bugs when given the diff version of the code and 0 out of 10 instances when presented with the raw source code. This result is weaker than it may appear at first glance. The defect in question involved an inverted authorization check in the code, allowing users to delete their accounts even if they were not authorized to do so. The AI reviewer ran the test on August 21, 2026, using a local model named local-advisor-v1:latest, which runs on Ollama with the Qwen35 architecture, 4.7 billion parameters, Q4_K_M quantization, and an 8,192-token context window. The AI reviewer instructed the tool to look for defects related to authorization and correctness. The harness used for the review requested a temperature of 0.1 and schema-constrained JSON output, with no random seed. Despite the AI reviewer's hand reading the defect in the diff version, none of the 20 valid runs resulted in a successful diagnosis by the tool itself. The tool requires the defect to be named in the findings, along with a file and line reference, and evidence copied from that line. In all valid runs, the findings list remained empty, as the model placed its diagnosis in the free-text control and evidence fields, which the tool labels as unverified model claims. One significant factor that may have contributed to the difference in results was the --trust-material-paths flag, which instructs the harness that the material is a diff and extracts file paths from it for scope checks. This flag enabled the model to score 9 of 10 diagnoses in the diff version, while the raw source code yielded no diagnoses in any of the 10 runs. The size of the input material also played a role, as the smallest and largest inputs both scored 0, while the middle-sized diff input scored 9 of 10. The AI reviewer recommends testing reviewers on a seeded defect first, comparing raw-source and diff inputs in actual review workflows, and reporting both the number of diagnoses among valid outputs and successful diagnoses across all attempts, while counting crashes as failures.",
  "summary": "Limits, before anything else. 21 attempts on one defect are repeated observations, not 21 independent tests. One seeded bug: an inverted authorization check in material I wrote myself. Run on 21 August 2026. One small local model. The alias local-advisor-v1:latest is local. Today ollama show reports architecture qwen35 , 4.7B parameters, Q4_K_M quantization, 8,192-token context, temperature 0.1.…",
  "key_points": [
    "AI reviewer identified 9 out of 10 bugs in diff version",
    "Failed to diagnose any defects in raw source code",
    "--trust-material-paths flag enabled diff version diagnoses"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}