arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37076cs.LG

UnlearningSoup:大型语言模型遗忘是否需要重复调参?

UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?

Puning Yang, Qizhou Wang, Junchi Yu, Bo Han, Xiuying Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM遗忘需重复调参的难题,提出UnlearningSoup框架,利用权重空间共享性能盆地,通过高效插值与重加权汤策略,实现2.4-3.3倍效率提升并改善性能。

中文摘要 AI 辅助

在广泛语料上训练的大型语言模型固有地存在记忆有害内容的风险,这些内容可能在后续输出中重新出现。为缓解这一问题,现有的遗忘方法通常依赖基于训练的参数更新,如梯度上升及其变体,以删除目标内容同时保留其他知识。然而,平衡遗忘与保留这两个相互竞争的目标使得这些方法的超参数选择尤为困难,往往需要重复调参才能获得强模型,且仍留有大量改进空间,并在不同模型和数据集间迁移性差。为应对这一挑战,我们探究遗忘运行是否在权重空间中展现出可利用的结构,并观察到来自不同运行的模型仍位于一个共享的评估性能盆地中。这表明,通过一种针对遗忘定制的汤(soup)策略,可以恢复出更强的模型,从而减少为进一步改进或新设置而进行重复调参的需求。受此启发,我们提出UnlearningSoup,一个统一框架,提供两种策略:EfficientSoup利用基于二分搜索的插值,在早期阶段快速发现性能良好的模型,而在此阶段重复调参会使强模型选择代价高昂。PerformanceSoup利用重新加权汤(souping)在后期阶段高效解锁剩余性能潜力,而此阶段重复调参变得日益低效。跨多种数据集和模型的广泛实验表明,UnlearningSoup在超参数选择上带来2.4倍至3.3倍的效率提升,同时在各设置下持续改善性能。

英文摘要

Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this issue, existing unlearning methods typically rely on training-based parameter updates, such as gradient ascent and its variants, to delete targeted content while preserving other knowledge. However, balancing the competing goals of forgetting and retention makes hyperparameter choices for these methods particularly difficult, often requiring repeated tuning to obtain a strong model that still leaves substantial room for improvement and transfers poorly across models and datasets. To address this challenge, we investigate whether unlearning runs exhibit exploitable structure in weight space, and observe that models from different runs still lie in a shared evaluation-performance basin. This suggests that stronger models may be recovered through an unlearning-tailored soup strategy, reducing the need for repeated tuning for further improvement or new settings. Motivated by this, we propose UnlearningSoup, a unified framework that provides two strategies: EfficientSoup uses binary-search-based interpolation to quickly discover a well-performing model in the early stage, where repeated tuning would otherwise make strong model selection costly. PerformanceSoup uses reweighted souping to efficiently unlock the remaining performance potential in the later stage, where repeated tuning becomes increasingly inefficient. Extensive experiments across diverse datasets and models show that UnlearningSoup delivers 2.4x to 3.3x efficiency gains in hyperparameter selection, while consistently improving performance across settings.

发表机构

  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)
  • RIKEN Center for Advanced Intelligence Project(理化学研究所先进智能研究中心)
  • Hong Kong Baptist University(香港浸会大学)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑