Beyond Marginal Coverage: Class-Conditional Conformal Prediction in Multi-Class Psychiatric Neuroimaging
Conformal prediction is increasingly proposed for clinical decision support because it provides distribution-free coverage guarantees that hold for any model and any data distribution. The guarantee routinely reported, however, is marginal, and we show that in multi-class diagnostic classification it can be satisfied exactly while individual diagnoses are covered very unevenly. On a four-way mood…
Conformal prediction is gaining traction for clinical decision support due to its ability to offer distribution-free coverage guarantees applicable to any model and data distribution. However, the commonly reported coverage guarantee is marginal, and the study reveals that in multi-class diagnostic classification, it can be met exactly while individual diagnoses experience highly uneven coverage.
Utilizing a four-way mood and psychosis classification task involving 1,520 subjects from three studies and 14 acquisition sites, the research demonstrates that marginal split-conformal calibration achieved an empirical coverage of 0.9000, compared to the target of 0.90. Yet, coverage varied significantly between groups - healthy controls at 0.941 and schizoaffective disorder at 0.818. The disparity between these figures is neutralized, rendering the overall result inconclusive.
The study further explores class-conditional (Mondrian) calibration, which mitigates the disparity in coverage by a mere 12.3 points, resulting in a minimal 0.4-point difference. However, this improvement comes at the expense of a 0.07 label reduction in mean set size. An intriguing consequence of maintaining consistent coverage across diagnoses is the inability to attribute residual variation in set size to class prior or per-class accuracy.
Instead, this variation becomes a distinct property of the subject. The prediction sets thus classify subjects into four categories: confident, boundary, ambiguous, and unresolved. The proportion of subjects flagged as label-ambiguous by a structural-MRI model trained separately on the same cohort increases gradually across these strata (34.1%, 57.3%, 68.1%, 81.8%; p = 8.8e-18).
Notably, no schizoaffective subject and only 0.9% of bipolar subjects reach the confident stratum, while 18.0% of controls and 13.4% of schizophrenia subjects do. Furthermore, set size at matched coverage provides a means to compare representations that accuracy alone cannot. When applied to structural MRI, functional MRI, and their fusion, the findings reveal a reversal in sign that accuracy fails to capture: fusion reduces set size for bipolar and schizoaffective subjects, while increasing it for controls.
The aggregate benefit of fusion changes sign below alpha = 0.10, while top-1 accuracy fluctuates from 0.495 to 0.578 to 0.611 across the three models. The researchers caution that reporting marginal coverage alone is insufficient when diagnostic classes are unbalanced.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.