pipette
ENEnglish

Privacy-Aware Distillation of Large Language Models for Enhanced Multimorbidity Scoring

R. Awasthi, Y. Yang, M. Li, X. Zhu

PreprintUso en el mundo real

En palabras de los autores

The truthful use of large language models (LLMs) is a growing challenge in safeguarding sensitive patient data from leakage. We introduce and evaluate a privacy-preserving knowledge distillation framework for LLM-based clinical modeling, using multimorbidity scoring as a healthcare task. Although LLMs can encode rich clinical knowledge and improve upon traditional rule-based comorbidity scoring, their direct evaluation on large-scale biobank data remains constrained by patient privacy. In our framework, multimorbidity reasoning is distilled from state-of-the-art LLM teacher models into compact student models (CoLLMs) using synthetic cohorts that preserve UK Biobank distributions, without exposing real patient data. This approach achieves high-fidelity knowledge transfer (Spearman {rho} = 0.75-0.89). Independent LLM-as-Judge evaluation confirms the clinical significance of the distilled knowledge and reveals substantial variability among teacher models. When applied to real UK Biobank data, CoLLM-derived multimorbidity scores improve survival prediction (C-index up to 0.91) and exhibit higher SNP heritability (h2 {approx} 0.05). Our work establishes a trustworthy, privacy-compliant pathway for large-scale healthcare applications of LLMs.

Resultado principalLimitación que admiten los autores

Apareció: martes, 22 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.

DOI: 10.64898/2026.09.19.26363476