arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

忠实双约束擦除用于鲁棒的大语言模型安全对齐

Faithful Dual-constrained Erasure for Robust LLM Safety Alignment

Jiaqing Li, Shide Zhou, Zhibo Zhang, Yuxi Li, Tianlong Yu, Kailong Wang

arXiv 2609.39279首次发表:更新:

发表机构

Hubei University(湖北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器遗忘后模型易受重训练攻击的脆弱性,提出双约束子空间投影框架FDCU,通过Fisher信息保留通用知识并禁止虚假抑制器激活,实现真正记忆删除,在保持通用性能的同时达到最先进的抗重训练鲁棒性。

AI 中文摘要

机器遗忘已成为从大语言模型(LLMs)中移除危险知识并实施安全对齐的关键机制。然而,近期研究揭示了一个持续存在的安全风险:经过遗忘处理的模型仍然极易受到重训练攻击,即在良性微调后,被抑制的恶意行为会迅速重新浮现。在本工作中,我们研究了遗忘的优化动态,并识别出这种脆弱性源于浅层对齐。模型并非有效擦除目标知识,而是常常利用一种捷径:激活先前休眠的参数作为虚假抑制器,在完整保留的恶意表征之上形成脆弱的抑制外壳。为解决此问题并强制实现真正的记忆删除,我们提出了FDCU,一种新颖的双约束子空间投影框架。FDCU通过一种高度可扩展的逐元素双掩蔽规则来限制参数更新:它通过Fisher信息保留通用知识流形,并通过最小功能干预原则(PMFI)严格禁止虚假抑制器的异常激活。通过可靠地阻断模型表面隐藏知识的能力,FDCU促进了目标表征的真正瓦解。在特定知识擦除和安全输出控制任务上的大量实验表明,FDCU在对抗重训练攻击方面达到了最先进的鲁棒性,同时保持了近乎无损的通用效用,确保了大语言模型的持久安全性。

英文摘要

Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models (LLMs). However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed malicious behaviors rapidly resurface after benign fine-tuning. In this work, we investigate the optimization dynamics of unlearning and identify that this vulnerability stems from shallow alignment. Rather than effectively erasing target knowledge, models often exploit a shortcut by activating previously dormant parameters to act as spurious suppressors, forming a fragile inhibitory shell over intact malicious representations. To address this issue and enforce authentic memory deletion, we propose FDCU, a novel dual-constrained subspace projection framework. FDCU restricts parameter updates through a highly scalable, element-wise dual-masking rule: it preserves general knowledge manifolds via Fisher Information and strictly prohibits the abnormal activation of spurious suppressors via the Principle of Minimal Functional Intervention (PMFI). By reliably blocking the model's ability to superficially hide knowledge, FDCU promotes the authentic dismantling of target representations. Extensive experiments across specific knowledge erasure and safe output control tasks demonstrate that FDCU achieves state-of-the-art robustness against retraining attacks while maintaining near-lossless general utility, ensuring durable safety for LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑