arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用关键神经元隔离剪枝防御大语言模型后门攻击

Defense Against LLM Backdoors using Critical Neuron Isolation Pruning

Yuxi Li, Zhibo Zhang, Kailong Wang, Xingshuo Han, Ling Shi, Haoyu Wang

arXiv 2607.19894首次发表:更新:

发表机构

Huazhong University of Science and Technology; Nanjing University of Aeronautics and Astronautics; Nanyang Technological University(华中科技大学; 南京航空航天大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型后门攻击问题,提出DeCNIP方法,利用表征分析统一识别和中和后门,通过优化交叉熵损失发现潜在威胁,隔离并剪枝后门关键神经元,有效降低攻击成功率,保持模型性能。

AI 中文摘要

大语言模型易受后门攻击,隐藏触发器会引发恶意输出。现有防御方法存在关键局限,如专注基于微调的后门,无法应对绕过训练管道的模型编辑攻击,且不适用于开放式生成场景。为弥补差距,引入DeCNIP,通过优化有害提示与候选令牌及良性输入间的交叉熵损失识别类似触发器行为,发现潜在威胁,隔离并选择性剪枝后门关键神经元以消除恶意影响。在六个开源大语言模型和两个基准数据集上的评估表明,DeCNIP将攻击成功率相对降低超95%,仅0.1%的神经元干预,同时保持正常基准下97%的模型性能,展现出有效性、鲁棒性和可扩展性。

英文摘要

Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or training-time mitigation, but face two key limitations. First, they focus on fine-tuning-based backdoors (e.g., PEFT modules) and fail to address insidious model-editing attacks that bypass training pipelines. Second, they target simple classification settings and do not naturally extend to open-ended LLM generation and do not naturally extend to the open-ended generation characteristics of LLMs. Consequently, these methods focus on surface-level behavioral patterns while neglecting the deeper representational causes of malicious activations. This lack of mechanistic understanding forces defenses to depend on empirical heuristics, limiting their robustness, generality, and practical applicability in real-world LLM deployment. To bridge this gap, we introduce DeCNIP (Defense with Critical Neuron Isolation Pruning), which leverages representational analysis to identify and neutralize backdoors in a unified pipeline. Specifically, DeCNIP identifies trigger-like behaviors by optimizing a cross-entropy loss between harmful prompts with candidate tokens and benign inputs. This representational discovery exposes latent threats by uncovering mechanisms through which triggers hijack model weights. It then isolates Backdoor Critical Neurons (BCNs) and prunes them selectively to remove malicious influence while preserving model utility. Extensive evaluations on six open-source LLMs and two benchmark datasets demonstrate that DeCNIP achieves over 95% relative reduction in Attack Success Rate (ASR), outperforming seven state-of-the-art defenses with only 0.1% neuron intervention. Moreover, it maintains 97% of the model's performance on normal benchmarks, demonstrating its efficacy, robustness, and scalability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑