arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

越受欢迎,越难遗忘:用于大语言模型遗忘的自适应受欢迎度

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

Anna Borisiuk, Andrey Savchenko, Alexander Panchenko, Elena Tutubalina

arXiv 2608.14229首次发表:更新:

发表机构

AIRI; Sber AI Lab; Skoltech; ISP RAS Research Center for Trusted Artificial Intelligence(AIRI; Sber AI实验室; 斯科尔科沃科学技术研究院; 俄罗斯科学院信息学研究所可信人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有LLM遗忘方法未区分训练数据频率的问题,提出AdaPop方法,经实验其在转述查询和对抗性重构下均能显著减少遗忘内容泄露。

AI 中文摘要

常见事实在预训练期间被更深地记忆,比罕见事实更难被移除,但现有的大语言模型(LLM)遗忘方法无论训练数据频率如何都采用统一的梯度压力。我们提出AdaPop(自适应受欢迎度)方法,该方法结合局部 token 置信度与从外部代理(如 Wikidata 站点链接、LLM 作为评判者)得出的每个事实依赖受欢迎度的指数,并通过双上升控制器自动调整每一轮的保留惩罚,以实现遗忘-保留平衡。在三个模型家族和两个基准测试中,在转述查询下,AdaPop 泄露的遗忘内容比对比方法少约5倍;在对抗性重构下,泄露量少约1.6倍。我们通过内部指标支持分析:与其他方法相比,我们的方法下,遗忘集的隐藏状态与遗忘前模型的状态距离更远,而保留集的表示仍保持接近。

英文摘要

Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM unlearning methods apply uniform gradient pressure regardless of training-data frequency. We propose the AdaPop (Adaptive Popularity) method, which combines local token confidence with a per-fact popularity-dependent exponent derived from an external proxy (e.g., Wikidata sitelinks, LLM-as-Judge), and automates the forget-retain balance via a dual-ascent controller that adjusts the retain penalty each epoch. Across three model families and two benchmarks, AdaPop leaks ~5x less forgotten content than competing methods under paraphrased queries and ~1.6x less under adversarial reformulations. We support our analysis with internal metrics: under our method, forget-set hidden states move further from the pre-unlearning model's states than under other methods, while retain-set representations remain close.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑