arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于持续大语言模型遗忘的轨迹引导遗忘-恢复网络

Trajectory-Guided Forget-Recover Network for Continual LLM Unlearning

Zezheng Wu, Xinghe Cheng, Qinggang Zhang, Haoran Luo, Jiapu Wang, Qing Yang, Jingwei Zhang

arXiv 2608.03123首次发表:更新:

AI 中文总结

针对持续大语言模型遗忘的两大挑战,提出TFR-Net,通过跟踪通道级风险抑制持久目标相关通道、恢复休眠通道,在四个数据集上取得更优的遗忘效果与保留性能权衡。

AI 中文摘要

机器遗忘旨在消除敏感数据对模型的影响。现实世界中,遗忘请求持续到来,引发两大挑战:其一,遗忘干预可能会将与目标相关的计算重新分配到剩余通路,致使先前遗忘的知识重新出现;其二,重复的遗忘干预可能会逐步降低模型保留实用性能所需的容量。为应对这些挑战,我们提出了轨迹引导遗忘-恢复网络(Trajectory-guided Forget-Recover Network,TFR-Net)。TFR-Net会跟踪各请求的通道级风险,将与目标相关的持久通道与瞬态热点分离,且仅抑制持久通道;同时,TFR-Net会通过重新激活休眠通道来恢复模型容量,这些通道对保留的实用性能有重要贡献,且当前和历史遗忘风险均较低,仅当保留的实用性能下降处于预定义容差范围内时,才会接受此类恢复。在四个数据集上的实验表明,TFR-Net在遗忘效果与保留实用性能之间始终能取得比代表性基准更优的权衡。

英文摘要

Machine unlearning aims to eliminate the influence of sensitive data on a model. In the real world, unlearning requests arrive continually, which gives rise to two challenges. First, an unlearning intervention may redistribute target-related computation across remaining pathways, allowing previously forgotten knowledge to re-emerge. Second, repeated unlearning interventions may progressively reduce the model capacity needed to preserve retained utility. To address these challenges, we propose the Trajectory-guided Forget-Recover Network (TFR-Net). TFR-Net tracks channel-level risk across requests. It separates persistent target-related channels from transient hotspots and suppresses only the persistent ones. TFR-Net also recovers model capacity by reactivating dormant channels. These channels make strong contributions to retained utility and show low current and historical forget risk. The recovery is accepted only when retained-utility degradation remains within a predefined tolerance. Experiments on four datasets show that TFR-Net consistently achieves a more favorable trade-off between unlearning effectiveness and retained utility than representative baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑