arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LACUNA: 评估大语言模型遗忘定位精度的测试平台

LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

Matteo Boglioni, Thibault Rousset, Siva Reddy, Marius Mosbach, Verna Dankers

arXiv 2607.02513首次发表:更新:

发表机构

Mila – Quebec Artificial Intelligence Institute; McGill University(米拉-魁北克人工智能研究所; 麦吉尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LACUNA测试平台,通过注入合成个人身份信息到模型参数,评估遗忘方法是否真正擦除知识,发现现有方法定位不精确且易受重现攻击,而精确定位可实现强擦除。

AI 中文摘要

大语言模型会记忆敏感训练数据,包括个人身份信息(PII),因此迫切需要可靠的后期移除方法。遗忘已成为一种有前景的解决方案,最先进(SOTA)方法通常遵循先定位、后遗忘的范式,针对特定模型参数。然而,现有基准仅在输出层面评估遗忘,留下了遗忘是否真正从模型参数中擦除知识或仅仅掩盖知识的疑问,而重现攻击的成功强化了这一担忧。为弥补这一差距,我们引入了LACUNA:首个具有真实参数级定位的遗忘测试平台。LACUNA通过掩码连续预训练将合成个体的PII注入到基于OLMo的1B和7B模型的预定义参数中,从而能够直接评估遗忘是否针对负责知识存储的权重。我们使用LACUNA对当前SOTA遗忘方法进行基准测试,发现尽管在输出层面表现强劲,现有方法高度不精确且易受重现攻击。我们进一步表明,当定位成功时,即使是简单的基于梯度的遗忘方法也能实现强擦除和对重现攻击的鲁棒性,突显了精确定位遗忘的重要性。我们发布LACUNA以补充行为评估,并推动基于定位的鲁棒遗忘的进一步进展。

英文摘要

LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods. Unlearning has emerged as a promising solution, with state-of-the-art (SOTA) methods often following a localize-first, unlearn-second paradigm that targets specific model parameters. However, existing benchmarks evaluate unlearning solely at the output level, leaving open the question of whether unlearning truly erases knowledge from a model's parameters or merely obfuscates it, a concern reinforced by the success of resurfacing attacks. To bridge this gap, we introduce LACUNA: the first unlearning testbed with ground-truth parameter-level localization. LACUNA injects PII of synthetic individuals into predefined parameters of 1B and 7B OLMo-based models via masked continual pretraining, enabling direct evaluation of whether unlearning targets the weights responsible for knowledge storage. We use LACUNA to benchmark current SOTA unlearning methods and find that, despite strong output-level performance, existing methods are highly imprecise and susceptible to resurfacing attacks. We further show that when localization is successful, even a simple gradient-based unlearning method achieves strong erasure and robustness to resurfacing attacks, highlighting the importance of precise unlearning. We release LACUNA to complement behavioral evaluations and drive further advances in robust, localization-based unlearning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑