Privacy-Aware Distillation of Large Language Models for Enhanced Multimorbidity Scoring
En palabras de los autores
The truthful use of large language models (LLMs) is a growing challenge in safeguarding sensitive patient data from leakage. We introduce and evaluate a privacy-preserving knowledge distillation framework for LLM-based clinical modeling, using multimorbidity scoring as a healthcare task. Although LLMs can encode rich clinical knowledge and improve upon traditional rule-based comorbidity scoring, their direct evaluation on large-scale biobank data remains constrained by patient privacy. In our framework, multimorbidity reasoning is distilled from state-of-the-art LLM teacher models into compact student models (CoLLMs) using synthetic cohorts that preserve UK Biobank distributions, without exposing real patient data. This approach achieves high-fidelity knowledge transfer (Spearman {rho} = 0.75-0.89). Independent LLM-as-Judge evaluation confirms the clinical significance of the distilled knowledge and reveals substantial variability among teacher models. When applied to real UK Biobank data, CoLLM-derived multimorbidity scores improve survival prediction (C-index up to 0.91) and exhibit higher SNP heritability (h2 {approx} 0.05). Our work establishes a trustworthy, privacy-compliant pathway for large-scale healthcare applications of LLMs.
Apareció: martes, 22 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.