RADNPO:用于大语言模型机器遗忘的无参考自适应负偏好优化
RADNPO: Reference-free Adaptive Negative Preference Optimization for LLM Unlearning
浏览论文内容
中文总结 AI 辅助
针对大语言模型遗忘中概率重分配失控的问题,提出无参考自适应负偏好优化(RADNPO),通过对比替代词元并自适应调整遗忘强度,在TOFU和MUSE上实现更优的遗忘质量与效用权衡。
中文摘要 AI 辅助
大型语言模型(LLMs)在预训练过程中可能会记忆敏感、私有或受版权保护的内容,这使得机器遗忘成为移除目标知识的必要手段。近年来,基于偏好优化(PO)的遗忘方法通过引入对齐风格的目标函数,比基于梯度上升(GA)的方法具有更好的稳定性,能有效抑制遗忘目标的概率。然而,仅抑制目标并不能充分约束遗忘后的下一个词元分布。现有方法对抑制后的概率质量如何重新分配的控制有限,并且未能根据目标置信度和分布集中度充分调整遗忘强度。即使在目标抑制之后,概率质量可能仍集中在少数非目标词元上,可能导致重复或无信息的输出。为解决这些局限性,我们提出了无参考自适应负偏好优化(RADNPO),该方法显式引导下一个词元概率的重新分配。具体而言,RADNPO将每个遗忘目标与当前下一个词元分布所偏好的替代词元进行对比,并利用目标置信度和下一个词元集中度自适应地调整词元级别的遗忘强度。在TOFU和MUSE上的实验表明,RADNPO在遗忘质量与模型效用之间取得了比现有基线更好的权衡。
英文摘要
Large language models (LLMs) can memorize sensitive, private, or copyrighted content during pre-training, making machine unlearning necessary for removing targeted knowledge. Recent preference optimization (PO)-based unlearning methods improve stability over gradient ascent (GA)-based methods by introducing alignment-style objectives, which effectively suppress the probability of forget targets. However, target suppression alone does not sufficiently constrain the next-token distribution after unlearning. Existing methods provide limited control over how suppressed probability mass is redistributed and insufficiently adapt forgetting strength to target confidence and distributional concentration. Even after target suppression, probability mass may remain concentrated on a few non-target tokens, potentially producing repetitive or uninformative outputs. To address these limitations, we propose Reference-free ADaptive Negative Preference Optimization (RADNPO), which explicitly guides next-token probability redistribution. Specifically, RADNPO contrasts each forget target with alternative tokens favored by the current next-token distribution and adaptively modulates token-level forgetting strength using target confidence and next-token concentration. Experiments on TOFU and MUSE demonstrate that RADNPO achieves a better trade-off between forgetting quality and model utility than current baselines.
发表机构
- Beihang University(北京航空航天大学)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。