{
  "id": 4550718,
  "title": "Same 8 drafts, one reviewer said revise 2, the other revise 7: calibrating rubrics for AI-on-AI review",
  "url": "https://urgent.news/2026/08/31/same-8-drafts-one-reviewer-said-revise-2-the-other-revise-7",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-31T02:17:00.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/rulestack/same-8-drafts-one-reviewer-said-revise-2-the-other-revise-7-calibrating-rubrics-for-ai-on-ai-2mp4"
  },
  "original_language": "en",
  "account": "Two AI reviewers were given the same eight reply drafts, rubric, and instructions. One reviewer found two drafts needing revision, while the other found seven. Initially, it was unclear whether one reviewer was flawed. However, this discrepancy led to a deeper understanding of how to calibrate rubrics for AI-on-AI review. The key causes of disagreement were unstated tolerance, batch-level rules being sensitive to batch size, and a poisoned premise. To address these issues, rubric items should now explicitly state tolerance, counting rules should be defined at the batch level, and premises should be labeled by verification recency. By treating disagreement as a signal, the focus should be on calibrating the rubric rather than the reviewers.",
  "summary": "We handed the same eight reply drafts, the same scoring rubric, and the same instructions to two independent AI reviewers. One returned revise 2 of 8 . The other returned revise 7 of 8 . If your first instinct is \"one of them is broken,\" it was ours too. It's also wrong, and the actual explanation reshaped how we write rubrics for any AI-on-AI review — code review, tone review, product QA, all of…",
  "key_points": [
    "Two AI reviewers gave conflicting feedback on eight drafts",
    "Discrepancy led to deeper understanding of rubric calibration",
    "Rubric items should state tolerance, rules defined at batch level"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}