Two identical runs scored 89 and 89. Two cases had flipped.
Read on: the A/B test this corrects · 繁體中文版 Two weeks ago I published an A/B test: I cut 41 AI tools' self-descriptions roughly in half, then ran a behavioural question bank against both versions to check that trigger rate hadn't dropped. Before: 88/96. After: 90/96. I wrote this sentence about it: ±2 cases at this sample size is noise, so I am not claiming it got better. A reader named Vinh…
Two AI tools each scored 89 out of 96 in an A/B test, where the original description was shortened by half. However, an A/B test by the author revealed two cases flipped between the two versions, cancelling each other out. The author discovered that the tool only showed failures, not lucky wins, leading to an incomplete understanding of the results.
By modifying the tool to record per-case vote counts, the author found that the majority of variance occurred in negative cases. A follow-up experiment confirmed that negative cases were more prone to flipping, while positive cases rarely experienced flips. The author concluded that churn is caused by negative cases wobbling between 1-of-3 and 2-of-3, while damage is indicated by positive cases collapsing to unanimous failure.
The total score alone cannot differentiate between churn and damage, emphasizing the importance of examining individual case votes.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.