pipette
ESEspañol

Leveraging Large Language Models for Temporal PHI De-identification in Real-World Clinical Notes

X. Wang, L. Wang, A. Wen, R. Li, S. Lu, Y. Hu, X. Li, H. Lyu, H. Liu

PreprintReal-world use

In the authors' words

De-identifying temporal protected health information (PHI) is challenging because temporal expressions appear in diverse formats and must be transformed while preserving clinically meaningful timelines. This study proposes and evaluates an LLM-based framework for temporal PHI de-identification in real-world clinical notes. Using 1,148 notes from 30 sarcoma patients (28,431 annotated temporal entities), we evaluated multiple modern LLMs for temporal entity extraction and surrogate generation. Proprietary models achieved strong extraction performance, with GPT-4o achieving the best overall results. Most models preserved temporal formatting (>99%), but surrogate generation remained challenging. GPT-5.4 achieved the highest shift correctness (90.2% on a shared-entity subset) and the lowest order violations. Error analysis revealed systematic shift deviations ({+/-}1, {+/-}30/31, {+/-}365 days), highlighting persistent limitations in LLM temporal reasoning. These findings suggest that reliable temporal de-identification will require hybrid or multi-agent approaches beyond standalone LLMs.

Main resultLimitation the authors admit

Appeared: Thursday, September 24. medRxiv. Preprint, not yet peer-reviewed.

DOI: 10.64898/2026.09.22.26360316