BLADE:用于大语言模型遗忘的双层低秩增广拉格朗日擦除方法
BLADE: Bilevel Low-rank Augmented-Lagrangian Erasure for LLM Unlearning
浏览论文内容
中文总结 AI 辅助
BLADE是一种受约束的双层框架,通过钳制熵遗忘损失、非对称增广拉格朗日和LoRA适配器双层结构解决LLM遗忘的鲁棒性问题,在三类基准上性能优于基线,且可应对规模扩大和连续遗忘步骤。
中文摘要 AI 辅助
现有大语言模型(LLM)遗忘方法存在鲁棒性问题:无约束的遗忘损失会降低模型连贯性,固定权重平衡无法在训练过程中保留数据难度变化时进行自适应调整,且在一个基准上有效的方法在规模扩大或重复应用时性能会下降。本文提出BLADE,一种受约束的双层框架,其三种机制可对优化空间进行平稳、可预测的控制:钳制熵遗忘损失,当token达到足够不确定性时梯度恰好为零;非对称增广拉格朗日,在任何违反保留要求后永久锁定保留保护;以及仅限于LoRA适配器的双层结构,在每个遗忘步骤前修复保留数据的损伤。该方法在三类基准系列中均表现优异,在TOFU上比最强基线的平均综合得分提高6%,在MUSE Books上提高9%,在KnowUndo上提高7%,且在MUSE News上进行4倍规模扩大和4次连续遗忘步骤时保持稳定,而最佳竞争方法则完全崩溃。
英文摘要
Existing LLM unlearning methods struggle with robustness: unbounded forget losses degrade model coherence, fixed-weight balancing cannot adapt as retain difficulty shifts mid-training, and methods that work on one benchmark falter under scaling or repeated application. We propose BLADE, a constrained bilevel framework whose three mechanisms give smooth, predictable control over the optimization landscape: a clamped-entropy forget loss whose gradient is exactly zero once a token reaches sufficient uncertainty; an asymmetric augmented Lagrangian that permanently ratchets retain protection after any violation; and a bilevel structure confined to LoRA adapters that repairs retain damage before each forgetting step. BLADE dominates across three benchmark families, improving average composite scores over the strongest baselines by $6$% on TOFU, $9$% on MUSE Books, and $7$% on KnowUndo, and it remains stable under $4\times$ scaling and $4$ sequential unlearning steps on MUSE News where the best competing method collapses entirely.
发表机构
- The Pennsylvania State University(宾夕法尼亚州立大学)
- American University(美利坚大学)
机构由 AI 辅助整理,请以论文原文为准。