pipette
ESEspañol

Discriminating rare disease cases from their controls based on observed and excluded phenotypes

Y. Guo, Y. Yao, Y. Zhang, G. Duan, N. Zhang, g. shan

PreprintReal-world use

In the authors' words

Rare diseases are individually uncommon but collectively prevalent. Their primary clinical challenge lies not in treatment but in diagnosis. In the early stages of clinical management, it is frequently unclear whether the observed phenotypes are associated with a rare disease. Leveraging machine learning methods to mine latent associations between these phenotypes and rare diseases for early diagnosis offers a viable strategy to alleviate this diagnostic dilemma. In the present study, we demonstrated that machine learning can effectively discriminate rare disease cases from their controls using observed and excluded phenotypes as features. Among them, the Random Forest model achieved the best classification performance with a certain degree of generalizability. Based on further analysis of the feature selection results, we conclude that the two factors, specificity and occurrence count, are important for phenotype selection in rare disease discrimination, and comparable importance should be attached to both observed and excluded phenotypes, during feature construction.

Main resultLimitation the authors admit

Appeared: Tuesday, September 22. bioRxiv. Preprint, not yet peer-reviewed.

DOI: 10.64898/2026.09.15.751721