Diagnostic Labels and Measurement Timing Drive Systematic Inconsistency in ADNI Neuroimaging Data
In the authors' words
The Alzheimer's Disease Neuroimaging Initiative (ADNI) is widely used to train machine learning models for Alzheimer's disease, yet whether it provides consistent ground truth for predictive modeling has not been systematically tested. In this paper, we showed that ADNI contains three interacting sources of bias with direct implications for machine learning: (a) diagnostic label inconsistency, (b) technical measurement drift, and (c) longitudinal survivor bias. A substantial proportion of cases, particularly within intermediate stages, fall outside ADNI's own diagnostic thresholds. MRI field strength and evolving processing pipelines introduce significant technical variability in hippocampal volume, while cohort survivor bias arising from differential retention of participants across phases further distorts longitudinal estimates of disease progression. These findings indicated that ADNI does not provide the stable, internally consistent labels often required in machine learning applications. We proposed a practical framework for diagnostic validation, feature harmonization, and cohort accounting, offering guidance for building more robust and biologically meaningful predictive models from large-scale neuroimaging cohorts.
Appeared: Wednesday, September 23. bioRxiv. Preprint, not yet peer-reviewed.