{
  "id": 6762119,
  "title": "Two identical runs scored 89 and 89. Two cases had flipped.",
  "url": "https://urgent.news/2026/09/11/two-identical-runs-scored-89-and-89-two-cases-had-flipped",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-11T14:05:10.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/dexterlung/two-identical-runs-scored-89-and-89-two-cases-had-flipped-3in3"
  },
  "original_language": "en",
  "account": "Two AI tools each scored 89 out of 96 in an A/B test, where the original description was shortened by half. However, an A/B test by the author revealed two cases flipped between the two versions, cancelling each other out. The author discovered that the tool only showed failures, not lucky wins, leading to an incomplete understanding of the results. By modifying the tool to record per-case vote counts, the author found that the majority of variance occurred in negative cases. A follow-up experiment confirmed that negative cases were more prone to flipping, while positive cases rarely experienced flips. The author concluded that churn is caused by negative cases wobbling between 1-of-3 and 2-of-3, while damage is indicated by positive cases collapsing to unanimous failure. The total score alone cannot differentiate between churn and damage, emphasizing the importance of examining individual case votes.",
  "summary": "Read on: the A/B test this corrects · 繁體中文版 Two weeks ago I published an A/B test: I cut 41 AI tools' self-descriptions roughly in half, then ran a behavioural question bank against both versions to check that trigger rate hadn't dropped. Before: 88/96. After: 90/96. I wrote this sentence about it: ±2 cases at this sample size is noise, so I am not claiming it got better. A reader named Vinh…",
  "key_points": [
    "Two AI tools scored 89 out of 96 in A/B test",
    "Author discovered two cases flipped between versions",
    "Negative cases more prone to flipping than positive cases"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}