arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

只遗忘重要的:面向鲁棒大语言模型的层选择性遗忘

Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

Ravi Ranjan, Olivera Kotevska, Agoritsa Polyzou

arXiv 2609.10439首次发表:更新:

AI 中文总结

提出FOM-UL层选择性遗忘框架,通过遗忘-保留显著性分数选择Transformer层进行定向更新,在TOFU等基准上提升遗忘效果、保持效用并增强量化鲁棒性。

AI 中文摘要

大型语言模型(LLMs)能够记忆并复现敏感、受版权保护或其他不受欢迎的训练内容,从而引发隐私、安全和监管方面的担忧。机器遗忘为完全重新训练提供了一种实用的替代方案,但许多现有方法采用广泛或固定的参数更新,这可能会降低模型效用,并且在部署变化(如训练后量化)下仍然脆弱,此时被遗忘的知识可能部分重新出现。我们提出了通过遗忘层实现只遗忘重要的(FOM-UL),这是一种层级别的遗忘框架,使用遗忘-保留显著性分数来选择Transformer层。该分数识别对遗忘集具有高影响力且对保留集低敏感性的层,使FOM-UL能够将更新集中在最有效的区域,同时保持模型大部分不变。这种有针对性的更新策略改善了遗忘-效用权衡,并通过减少小而分散的更新被低位舍入消除的机会,为量化鲁棒的遗忘提供了一条经验路径。在TOFU、KnowUnDo和MUSE风格的评估中,与强基线GA、NPO、KLD、SURE、ReLearn和LUNAR相比,FOM-UL减少了残余记忆,同时将保留集效用保持在接近原始模型的水平。在8位和4位训练后量化下,FOM-UL比竞争方法保持了更强的记忆抑制和效用保留,对抗性提示评估显示被遗忘内容的恢复率更低。总体而言,FOM-UL提供了一种高效的遗忘策略,改善了目标遗忘、效用保留和部署鲁棒性,但不声称提供正式的擦除保证。

英文摘要

Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable training content, creating privacy, safety, and regulatory concerns. Machine unlearning offers a practical alternative to full retraining, but many existing methods apply broad or fixed parameter updates that can degrade utility and remain brittle under deployment changes such as post-training quantization, where forgotten knowledge may partially re-emerge. We propose Forgetting Only What Matters via Unlearning Layers (FOM-UL), a layer-level unlearning framework that selects transformer layers using a forget-to-retain significance score. This score identifies layers with high influence on the forget set and low sensitivity to the retain set, allowing FOM-UL to concentrate updates where they are most effective while leaving most of the model unchanged. This targeted update strategy improves the forgetting-utility trade-off and provides an empirical path toward quantization-resilient unlearning by reducing the chance that small, diffuse updates are erased by low-bit rounding. Across TOFU, KnowUnDo, and MUSE-style evaluations, FOM-UL reduces residual memorization compared with strong GA, NPO, KLD, SURE, ReLearn, and LUNAR-based baselines while preserving retain-set utility close to the vanilla model. Under 8-bit and 4-bit post-training quantization, FOM-UL maintains stronger memorization suppression and utility preservation than competing methods, and adversarial prompt evaluations show lower recovery of forgotten content. Overall, FOM-UL provides an efficient unlearning strategy that improves targeted forgetting, utility preservation, and deployment robustness without claiming formal guarantees of erasure.

Comments22 pages, 6 figures, 11 tables, AACL-IJCNLP 2026, conference paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑