pipette
ENEnglish

Machine Unlearning for Large Language Models: Foundations, Advances, and Agentic Extensions

Xiaoyu Xu, Minxin Du, Li Bai, Junxu Liu, Yaxin Xiao, Kun Fang, Liu Yang, Huadi Zheng, Peizhao Hu, Qingqing Ye, and Haibo Hu

Preprint

En palabras de los autores

Machine unlearning aims to remove target influence while preserving other capabilities. This survey compares methods, benchmarks, and evidence across large language models and systems using retrieval, memory, tools, and interacting agents. A five-layer framework connects removal requests, system boundaries, target locations, interventions, and supported claims. A seven-stage lifecycle and six evidence dimensions guide comparison. The review shows that target construction, retained data, and recovery tests affect reported outcomes. Evidence from model evaluations remains insufficient to establish removal across external state and subsequent updates, motivating evaluation that tracks dependencies and tests whether target influence returns.

Resultado principalLimitación que admiten los autores

Apareció: lunes, 28 de septiembre. arXiv. Preprint, todavía sin revisión por pares.

Comentario de los autores: Under submission