The Average Is Nobody's Result
In 2013 four researchers went back to a completed mammography study and asked it a question it had not been designed to answer. The original study had put 50 radiologists in front of 180 mammograms, twice. Once unaided, once with computer-aided detection marking suspicious regions. The finding was a null. On average, computer aid changed nothing measurable, and the profession moved on. Andrey…
In 2013, a group of researchers revisited a mammography study and asked a question it wasn't designed to answer. The original study involved 50 radiologists examining 180 mammograms, both with and without computer-aided detection (CAD). The result was a null finding, with computer aid not showing any measurable difference. Andrey Povyakalo and his colleagues at City University London reanalyzed the data by separating the readers instead of pooling them.
They discovered that CAD had two large effects, one increasing sensitivity for the least discriminating radiologists and another decreasing sensitivity for the most discriminating radiologists. The average effect canceled out, and the study reported no significant average effect. However, this average masked two opposite effects, one benefiting the less discriminating readers and the other harming the more discriminating readers.
This phenomenon is common in averages, where a mixture of effects can result in a seemingly neutral outcome even when significant differences exist. In other scenarios, such as measuring review throughput, tickets resolved, or defect-escape rates, the outcome can be positive while having a negative effect on a third of the team.
The nine percent improvement reported in a tool's performance is a real answer to a real question but cannot provide guidance on how to deploy the tool. The nine percent improvement could be attributed to various groups of personnel, each with its own set of requirements. Averages, like those presented in reports, can be misleading as they fail to capture the nuances of individual performance, potentially leading to incorrect decisions based on a single, inaccurate number.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written; read the original for the full account.



