arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Nullify:用于无需训练的大语言模型遗忘的零空间激活引导

Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning

Wei Zhai, Xiang Liu, Qiang Huang, Rui Qian, Lemao Liu, Ziwei Li, Ziqi Wang, Zhitao Huang, Dejing Dou

arXiv 2610.10655首次发表:更新:

发表机构

College of Computer Science and Artificial Intelligence, Fudan University; BEDI Cloud; School of Electrical and Information Engineering, Tianjin University; King Abdullah University of Science and Technology (KAUST); School of Software Technology, Zhejiang University; School of Transnational Law, Peking University(复旦大学计算机与人工智能学院; BEDI云公司; 天津大学电气与信息工程学院; 阿卜杜拉国王科技大学; 浙江大学软件学院; 北京大学国际法学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出无需训练的LLM遗忘方法Nullify,通过推理时激活引导平衡遗忘质量与模型效用,在TOFU、MUSE上表现优于或相当基准,且高效即插即用。

AI 中文摘要

大语言模型(LLMs)在预训练过程中不可避免地会内化大量敏感或私人信息,而LLM遗忘旨在选择性擦除特定知识以防止隐私泄露,同时最小化模型效用损失。然而,现有方法难以在遗忘质量与效用之间取得平衡,且通常因参数微调而产生大量计算成本。为解决这一问题,我们提出Nullify,一种用于LLM遗忘的无需训练、非破坏性的激活引导方法。Nullify在推理阶段使用引导向量,将与隐私相关的激活重定向至远离其记忆答案的方向,同时满足零空间约束,使保留查询的激活基本不受影响,以维持模型效用。在TOFU和MUSE上的评估显示,Nullify在遗忘质量上与既定基准相当或超越,同时实现了近乎无损的模型效用保留。由于完全避免了权重更新,Nullify是一种高效、即插即用的推理时干预框架。

英文摘要

Large Language Models (LLMs) inevitably internalize substantial amounts of sensitive or private information during pre-training, while LLM unlearning aims to selectively erase specific knowledge to prevent privacy leakage with minimal loss of model utility. However, existing methods struggle to balance forget quality with utility, and typically incur substantial computational costs due to parameter fine-tuning. To address this, we propose Nullify, a training-free, non-destructive activation steering method for LLM unlearning. Nullify employs steering vectors during inference to redirect privacy-related activations away from their memorized answers, while satisfying a null-space constraint that leaves retained-query activations essentially unaffected to maintain utility. Evaluations on TOFU and MUSE show that Nullify matches or surpasses established baselines in forget quality while achieving near-lossless preservation of model utility. By avoiding weight updates entirely, Nullify serves as an efficient, plug-and-play inference-time intervention framework.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑