Discriminating rare disease cases from their controls based on observed and excluded phenotypes
Rare diseases are individually uncommon but collectively prevalent. Their primary clinical challenge lies not in treatment but in diagnosis. In the early stages of clinical management, it is frequently unclear whether the observed phenotypes are associated with a rare disease. Leveraging machine learning methods to mine latent associations between these phenotypes and rare diseases for early…
Rare diseases are uncommon individually but collectively prevalent, presenting a primary clinical challenge not in treatment but in diagnosis. At the early stages of clinical management, it is often uncertain whether observed phenotypes are associated with a rare disease. Utilizing machine learning methods to uncover latent associations between these phenotypes and rare diseases for early diagnosis presents a viable solution to address this diagnostic challenge.
This study demonstrates that machine learning can effectively differentiate rare disease cases from their controls using both observed and excluded phenotypes as features. Among the models tested, the Random Forest algorithm demonstrated the best classification performance, exhibiting a certain level of generalizability. Further analysis of feature selection results revealed that two factors - specificity and occurrence count - are crucial for phenotype selection in rare disease discrimination.
Importantly, both observed and excluded phenotypes should be given comparable importance during the feature construction process.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.