arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAUL:感知锐度的增广拉格朗日模型遗忘算法

SAUL: Sharpness-Aware Augmented-Lagrangian Unlearning

Jaewan Choi, Junyoung Yang, Sangdon Park

arXiv 2608.16249首次发表:更新:

发表机构

POSTECH; Computer Science and Engineering, POSTECH; Graduate School of Artificial Intelligence, POSTECH(浦项科技大学; 浦项科技大学计算机科学与工程学院; 浦项科技大学人工智能研究生院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SAUL算法,将模型遗忘建模为显式约束的受约束最小化问题,通过增广拉格朗日控制器等实现更优遗忘-效用权衡,即插即用修改器可提升基线模型遗忘后效用。

AI 中文摘要

大语言模型(LLMs)的机器遗忘面临着擦除目标知识与保留通用效用之间的关键权衡难题。本文提出SAUL(Sharpness-Aware Augmented-Lagrangian Unlearning,感知锐度的增广拉格朗日模型遗忘算法),将模型遗忘问题建模为受约束最小化问题,遵循“遗忘充分但不过度”的原则。SAUL的核心在于将遗忘设定为具有明确满足准则的显式约束,而现有遗忘方法通常通过优化目标隐式指定所需的遗忘程度。增广拉格朗日控制器会根据约束违反情况自适应调整遗忘侧压力,当预设准则被满足时,最终可停用遗忘侧更新。对保留和遗忘目标的感知锐度更新,以及维持角色分离状态的双优化器设计,进一步稳定了遗忘动态过程。我们在TOFU、WMDP和MUSE基准上对SAUL进行评估,结果表明,在基准特定的遗忘准则下,与代表性的感知锐度类和扰动类基线相比,SAUL实现了更优的遗忘-效用权衡。除完整的SAUL框架外,我们还在TOFU数据集上验证,将增广拉格朗日控制器作为即插即用修改器应用于代表性基线,可提升其遗忘后的效用,凸显了显式遗忘控制的实用价值。

英文摘要

Machine unlearning in Large Language Models (LLMs) faces a critical trade-off between erasing target knowledge and preserving general utility. We propose SAUL (Sharpness-Aware Augmented-Lagrangian Unlearning), which formulates unlearning as a constrained minimization problem following the principle of "forget enough, but no more than necessary." At its core, SAUL formulates forgetting as an explicit constraint with a prescribed satisfaction criterion, whereas prior unlearning methods typically specify the desired level of forgetting implicitly through optimization objectives. An augmented Lagrangian controller adaptively adjusts forget-side pressure according to constraint violation and can eventually deactivate the forget-side update as the prescribed criterion remains satisfied. Sharpness-aware updates on both retain and forget objectives, together with a dual-optimizer design that maintains role-separated states, further stabilize the resulting unlearning dynamics. We evaluate SAUL on the TOFU, WMDP, and MUSE benchmarks, demonstrating favorable forgetting-utility trade-offs over representative sharpness- and perturbation-based baselines under benchmark-specific forgetting criteria. Beyond the complete SAUL framework, we further show on TOFU that applying the augmented-Lagrangian controller as a drop-in modifier to representative baselines improves their post-forgetting utility, demonstrating the practical value of explicit forgetting control.

Comments9 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑