arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨越边际悬崖:通过边际校准实现可再学习鲁棒的大语言模型遗忘

Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration

Xiangyu Yin, Jiaxu Liu, Zhen Chen, Chih-Hong Cheng

arXiv 2607.27836首次发表:更新:

AI 中文总结

该研究针对大语言模型遗忘在再学习攻击下的脆弱性,提出边际校准方法,在多个数据集和模型上提升遗忘鲁棒性并降低成员推断风险,仅以保留侧效用降低为代价。

AI 中文摘要

大语言模型遗忘在再学习攻击下始终存在脆弱性。在TOFU数据集上,对20个遗忘样本进行微调可大幅恢复所有评估方法的保留遗忘集ROUGE分数,我们将这种脆弱性追溯至优化几何。涵盖梯度、偏好和蒸馏类别的14种事后方法的每token答案边际,在42个方法-规模单元中的41个里收敛至保留参考值上方的窄带,我们将这种规律性称为“边际悬崖”。我们证明,当保留耦合将遗忘内容的诊断对数几率维持在阈值以上时,就会出现这种悬崖,而token饱和损失在平稳状态下会引发该条件,我们在42个单元中的34个直接验证了这一点。边际校准(Margin Calibration, MC)是一种即插即用的优化方法,它添加了一个以参考每token边际为锚点的非饱和边际 hinge,以及在不相交指令语料库上的KL探测,从而在原生损失饱和处恢复遗忘侧压力。在规定的梯度主导条件下(我们通过监测该优化方法测量其轨迹上的梯度特征),其平稳集位于悬崖跨越侧,为再学习边际提升提供了攻击预算上限。在TOFU(三种Llama-3规模、三种遗忘层级)、Llama-2-7B-hf上的MUSE-News以及Phi-3.5面板上,单一冻结配置在所有14项直接遗忘聚合和所有已填充再学习单元中均获胜(面板平均攻击后ROUGE-L从0.41降至0.18),并在14项中的13项降低了原始成员AUC,主要代价是保留侧效用降低。一种部署变体无需保留训练参考即可匹配这些增益。

英文摘要

Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tuning on twenty forget examples substantially recovers held-out forget-set ROUGE for every method we evaluate, and we trace this fragility to optimization geometry. The per-token answer margin of fourteen post-hoc methods spanning gradient, preference, and distillation families converges into a narrow band above the retain reference in 41 of 42 method--size cells, a regularity we call the margin cliff. We prove that this cliff follows whenever the retain coupling holds the diagnostic log-odds of forget content above a floor, a condition that token-saturating losses induce at stationarity and that we verify directly on 34 of 42 cells. Margin Calibration (\textsc{MC}) is a plug-in polish adding a non-saturating margin hinge anchored at the reference's per-token margin plus a KL probe on a disjoint instruction corpus, restoring forget-side pressure where the native loss saturates. Under a stated gradient-dominance condition, whose on-trajectory gradient signature we measure by instrumenting the polish, its stationary set lies on the cliff-crossing side, yielding an attack-budget upper bound on the relearn margin lift. Across TOFU (three Llama-3 sizes, three forget tiers), MUSE-News on Llama-2-7B-hf, and a Phi-3.5 panel, a single frozen configuration wins all 14 head-to-head forget aggregates and all populated relearn cells (panel-mean post-attack ROUGE-L $0.41$ to $0.18$) and lowers raw membership AUC on 13/14, with reduced retain-side utility as the main cost. A deployment variant matches these gains without a retain-trained reference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑