A neurocognitive speech taxonomy for voice biomarkers of Alzheimer's disease
In the authors' words
INTRODUCTION: Voice data combined with large language models may detect cognitive impairment, yet an interpretable framework has not been formally established. We developed and validated a scalable and interpretable framework, the Neurocognitive Speech Taxonomy (NST). METHODS: NST maps 314 voice features to 7 neurocognitive domains derived from neuropsychology, speech pathology, neurology and related fields. The domains capture features including articulatory precision, cognitive-linguistic, executive fluency & planning, phonation & laryngeal control, prosodic modulation, lexical-semantic, and morphosyntactic complexity. We evaluated NST structural stability, clinical validity and racial-educational equity in three cohorts (total N = 1,479). Cohort data included demographic and neuropsychological evaluations and AD biomarker status. We tested whether NST scores distinguish cognitive status (unimpaired vs impaired) and AD-biomarker status, track cognitive change, and correlate with AD neuroimaging measures. RESULTS: NST showed stable cross-cohort structure (Mantel r = 0.71-0.77) and differentiated MCI (d = -0.39) and AD biomarker status (d = -0.23). NST tracked longitudinal cognitive change, correlated with hippocampal volume (r up to 0.28, P < .001), and showed smaller racial and educational disparities relative to standard cognitive testing (NST d = 0.03-0.33, MoCA d = 0.37-0.72). DISCUSSION: NST provides a validated neurocognitive construct that facilitates interpretation of AI-driven voice biomarkers of AD. This study was not registered in a public trials registry.
Appeared: Saturday, September 26. medRxiv. Preprint, not yet peer-reviewed.