发表机构
Tsinghua University; Ant Group(清华大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出新基准unlearning评估LLM遗忘的鲁棒性,实验发现现有遗忘方法易受多跳推理路径和恢复攻击影响,还探究了遗忘质量、鲁棒性与模型效用的权衡。
AI 中文摘要
机器学习遗忘方法的基准测试对于理解大语言模型(LLMs)是否已移除敏感知识至关重要。当前的遗忘基准主要包含单跳问题和有限的多跳问题,虽有效但仍面临两大挑战:一是知识并非孤立存在,与普通查询相比,多样的多跳推理路径更可能引发知识泄露;二是遗忘可能存在脆弱性,被遗忘的知识可通过轻量级遗忘后适配等恢复攻击部分恢复,使得静态评估不足。因此,本文提出unlearning作为一种新基准,以理解LLM在多样推理路径和恢复攻击下的鲁棒知识移除。我们在3个模型、6种遗忘方法和2个精心整理的数据集上开展实验,结果显示现有方法易受多跳推理路径和恢复攻击影响。我们进一步探究了LLM遗忘中遗忘质量、鲁棒性与模型效用之间的权衡关系。
英文摘要
Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Current unlearning benchmarks include mainly single-hop questions and a narrow set of multi-hop questions. Although effective, they still face two challenges. (1) Knowledge is not isolated, whereby diverse multi-hop reasoning paths can potentially induce knowledge leakage than normal queries. (2) Unlearning may be fragile: unlearned knowledge can be partially recovered through recovery attacks such as lightweight post-unlearning adaptation, making static evaluation insufficient. Therefore, in this paper, we introduce \unlearning as a novel benchmark to understand robust LLM knowledge removal across diverse reasoning paths and recovery attacks. We experiment with this benchmark on 3 models, 6 unlearning methods, and 2 carefully curated datasets. Results show that existing methods are vulnerable to multi-hop reasoning paths and recovery attacks. We further explore the trade-off among forget quality, robustness, and model utility for LLM unlearning.
Comments19 pages, 7 figures