arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

移除大海捞针:通过权重正交化在大型语言模型中去除后门

Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation

Minoo Kim, Vasileios Lampos, George Drayson

arXiv 2610.00348首次发表:更新:

发表机构

Locai Labs; UCL(Locai实验室; 伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大型语言模型的后门攻击,提出无需训练的NEEDLE方法,通过权重正交化精准移除后门,保持模型性能与安全,实现最低攻击成功率。

AI 中文摘要

后门攻击可以在大型语言模型(LLMs)训练期间植入,当输入中出现触发器时会导致不良行为。现有的针对LLMs的后门防御方法试图移除后门,但无意中改变了模型对良性提示的输出分布,这可能导致模型性能和安全性下降。我们提出NEEDLE,一种无需训练的目标后门移除方法。一旦识别出触发器,我们的方法通过激活向量估计后门方向和拒绝子空间,然后应用顺序权重正交化来抑制后门,同时防止与拒绝相关的表示发生变化。NEEDLE既不需要干净的参考模型,也不需要原始的污染训练数据。我们在多个模型家族和攻击类型上进行了评估。NEEDLE在评估的防御方法中实现了最低的平均攻击成功率(ASR),包括在具有挑战性的代码注入攻击上达到0%,同时实现了最低的KL散度以及最小的能力和安全性变化。

英文摘要

Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original poisoned training data. Evaluation is conducted across multiple model families and attack types. NEEDLE achieves the lowest mean Attack Success Rate (ASR) among the evaluated defences, including 0% on challenging code injection attacks, while resulting in the lowest KL divergence and minimal changes in capability and safety.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑