arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TULIP:在逐输入识别的层上进行目标LLM遗忘

TULIP: Targeted LLM Unlearning at Layers Identified Per-Input

Yejin Kim, William F. Shen, Seokwon Jung, Daeun Park, Seong Joon Oh

arXiv 2609.34591首次发表:更新:

发表机构

Korea Advanced Institute of Science & Technology (KAIST); University of Cambridge; Sookmyung Women’s University(韩国科学技术院(KAIST); 剑桥大学; 淑明女子大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM遗忘中固定层干预的不足,提出TULIP方法,通过逐输入识别形成-读出边界并在该层移除隐藏状态对齐,在多个基准和模型上优于现有方法。

AI 中文摘要

表示级遗忘干预了LLM的中间隐藏状态。尽管知识分布在各个层中,现有方法对整个遗忘集在单一固定层上操作。我们质疑这样的固定层是否足够。为了回答这个问题,我们设计了一个劫持实验,将目标模型的隐藏状态移植到仅基于保留集训练的神谕模型中。神谕模型本身无法产生遗忘答案,但能从移植的状态中产生答案。因此,答案是在中间层形成的,之后仅被读出,所以遗忘应聚焦于形成而非读出。此外,形成结束的层因输入而异。受这些发现启发,我们提出了逐输入识别层的目标遗忘(TULIP)。对于每个输入,TULIP使用logit lens定位形成-读出边界,并在此处移除隐藏状态与遗忘答案的反嵌入向量的对齐。TULIP在TOFU、PISTOL和WMDP上,跨Llama、Qwen和Zephyr模型,始终优于输出级和表示级基线。它还对释义和量化攻击保持鲁棒性。除了独立使用外,其逐输入层选择作为即插即用组件,进一步改进了现有方法。

英文摘要

Representation-level unlearning intervenes on the intermediate hidden states of LLMs. Although knowledge is distributed across layers, existing methods operate at a single fixed layer for the entire forget set. We ask whether such a fixed layer is sufficient. To answer this, we design a hijacking experiment that grafts hidden states of the target model into an oracle trained only on the retain set. The oracle cannot produce the forget answer on its own, yet it produces the answer from the grafted state. Thus, the answer is formed at an intermediate layer and merely read out afterward, so unlearning should focus on formation, not readout. Moreover, the layer where formation ends varies widely across inputs. Motivated by these findings, we propose Targeted Unlearning at Layers Identified Per-input (TULIP). For each input, TULIP uses the logit lens to locate the formation-readout boundary and removes the hidden state's alignment with the forget answer's unembedding vector there. TULIP consistently outperforms output- and representation-level baselines on TOFU, PISTOL, and WMDP across Llama, Qwen, and Zephyr models. It also remains robust to paraphrase and quantization attacks. Beyond standalone use, its per-input layer selection serves as a plug-and-play component that further improves existing methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑