arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LoRA微调大语言模型的后门净化:基于零空间投影

Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection

Jianwei Li, Jung-Eun Kim

arXiv 2610.00685首次发表:更新:

发表机构

North Carolina State University(北卡罗来纳州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出零空间投影方法,无需触发器先验或重训练即可净化LoRA微调模型的后门,将攻击成功率从近100%降至10%以下,同时保留基础能力与下游技能。

AI 中文摘要

随着大语言模型(LLMs)和参数高效微调(PEFT)方法的快速普及,后门攻击的风险变得更加严重。现有的后门净化方法通常至少依赖于以下强假设之一:对触发器(trigger)的先验知识、访问干净参考数据,或进行激进的重新训练,并且它们往往缺乏全面的评估。这些限制极大地削弱了其实际适用性。为克服这些挑战,我们的工作提出在不依赖这些假设、甚至无需对可疑参数进行事后重新训练的情况下,净化LoRA微调的大语言模型。我们的目标是显著降低攻击成功率(ASR),同时保留(i)基础模型的通用能力以及(ii)通过适配器(adapter)学习到的新下游技能。通过一系列消融研究,我们逐步将方法从文本分类设置中的单层扩展到生成任务中的全参数大语言模型。通过精心设计的数据整理和特征近似,我们提取高保真的后门方向,并为每个层或头在输入和输出通道中构建正交零空间,LoRA更新被投影到这些零空间上。实验表明,我们的零空间投影方法将攻击成功率从接近100%降至10%以下,同时在下游任务适配过程中保留了基础模型的良性性能和适配器学到的能力。

英文摘要

With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model's general capabilities and (ii) the new downstream skills learned through the adapter. Through a series of ablation studies, we progressively scale our approach from a single layer in a text classification setting to a full-parameter LLM in the generative task. Through careful data curation and feature approximation, we extract high-fidelity backdoor directions and, for each layer or head, construct orthogonal null spaces in both the input and output channels, onto which the LoRA updates are projected. Empirically, our null-space projection method reduces the ASR from nearly 100% to less than 10%, while preserving the base model's benign performance and the adapter's learned abilities during downstream task adaptation.

CommentsNeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑