pipette
ESEspañol

Privacy-Aware Distillation of Large Language Models for Enhanced Multimorbidity Scoring

R. Awasthi, Y. Yang, M. Li, X. Zhu

PreprintReal-world use

In the authors' words

The truthful use of large language models (LLMs) is a growing challenge in safeguarding sensitive patient data from leakage. We introduce and evaluate a privacy-preserving knowledge distillation framework for LLM-based clinical modeling, using multimorbidity scoring as a healthcare task. Although LLMs can encode rich clinical knowledge and improve upon traditional rule-based comorbidity scoring, their direct evaluation on large-scale biobank data remains constrained by patient privacy. In our framework, multimorbidity reasoning is distilled from state-of-the-art LLM teacher models into compact student models (CoLLMs) using synthetic cohorts that preserve UK Biobank distributions, without exposing real patient data. This approach achieves high-fidelity knowledge transfer (Spearman {rho} = 0.75-0.89). Independent LLM-as-Judge evaluation confirms the clinical significance of the distilled knowledge and reveals substantial variability among teacher models. When applied to real UK Biobank data, CoLLM-derived multimorbidity scores improve survival prediction (C-index up to 0.91) and exhibit higher SNP heritability (h2 {approx} 0.05). Our work establishes a trustworthy, privacy-compliant pathway for large-scale healthcare applications of LLMs.

Main resultLimitation the authors admit

Appeared: Tuesday, September 22. medRxiv. Preprint, not yet peer-reviewed.

DOI: 10.64898/2026.09.19.26363476