发表机构
Manipal Institute of Technology, Manipal Academy of Higher Education; San José State University(马尼帕尔理工学院,马尼帕尔高等教育学院; 圣何塞州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
综述聚焦于网络防御中LLM遗忘问题,探讨其面临风险,核心问题是现有方法能否真的去除知识。主要关注基于梯度的方法,因其与现有训练管道兼容且可扩展,从多方面审视LLM遗忘,为解决相关问题提供参考。
AI 中文摘要
大语言模型(LLMs)越来越多地部署在医疗、金融、教育和决策支持等安全关键系统中,但它们无法遗忘会带来严重的网络安全、隐私和安全风险。敏感个人信息等在模型部署后仍编码在数十亿参数中,使模型易受攻击。由于重新训练数十亿参数模型计算上不可行,LLM遗忘成为主要网络防御措施。当前方法是真的去除了知识,还是仅阻止模型在普通提示条件下表达它,这一核心问题仍未解决。本综述从安全、鲁棒性和可验证遗忘角度审视LLM遗忘,主要关注基于梯度的方法,因其与现有训练管道兼容且可扩展到数十亿参数模型而在该领域占主导地位。
英文摘要
LLMs are increasingly deployed in security-critical systems across healthcare, finance, education, and decision support, yet their inability to forget creates serious cybersecurity, privacy, and safety risks. Sensitive personal information, copyrighted material, hazardous domain knowledge, and memorized training data remain encoded across billions of parameters long after deployment, leaving models vulnerable to extraction, jailbreak attacks, membership inference, and regulatory non-compliance. Real-world incidents, from chatbots regenerating private information to fabricated legal citations producing direct legal and financial cost, place the problem at the center of the emerging-threats landscape rather than the realm of speculation. Because retraining billion-parameter models on revised corpora is computationally infeasible, and because knowledge within an LLM is distributed and entangled across parameters rather than localized to identifiable units, LLM unlearning has emerged as the principal cyber defense response, aiming to remove or suppress targeted knowledge from a trained model without retraining and without eroding what the model should still know. A central question, however, remains unresolved. Do current methods genuinely remove knowledge, or do they only stop the model from expressing it under ordinary prompting conditions? This survey examines LLM unlearning through the lens of security, robustness, and verifiable forgetting, with primary focus on gradient-based methods, which have come to dominate the field due to their compatibility with existing training pipelines and their scalability to billion-parameter models.
Comments42 pages, 5 tables