发表机构
University of Louisville; University of North Texas(路易斯维尔大学; 北得克萨斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NeuronGuard是一种微调阶段防御方法,通过重新分配安全信号抵御越狱和神经元级攻击,在三个LLM、六种攻击下实现近乎零攻击成功率且保持任务准确率。
AI 中文摘要
大语言模型(LLM)的安全对齐对日益增多的攻击仍较为脆弱。越狱攻击通过精心设计的提示绕过安全机制,而神经元级攻击则在部署后直接剪枝安全关键神经元。两者均利用了一个共同弱点:与安全相关的信息集中在稀疏的神经元子集里。我们提出NeuronGuard,一种微调阶段的防御方法,通过在更广泛的神经元集合中重新分配安全信号,同时增强LLM对两类攻击的抵御能力。NeuronGuard通过定期更新的每层线性分类器动态识别安全关键神经元,在刻意的神经元消融下强制拒绝行为,并应用KL散度正则化实现分布一致性。随机梯度投影策略通过解决防御与任务目标间的冲突,保留下游任务效用。我们提供了NeuronGuard严格降低攻击成功率(ASR)上界的形式化保证,在三个LLM、六种最先进攻击策略及多模态设置下的实验证实,其在保持任务准确率的同时实现了近乎零的ASR,包括对抗白盒自适应攻击者的情况。
英文摘要
Safety alignment in large language models (LLMs) remains brittle against a growing spectrum of attacks. Jailbreak attacks bypass safety mechanisms through crafted prompts, while neuron-level attacks directly prune safety-critical neurons post-deployment. Both exploit a common weakness: safety-relevant information concentrates in a sparse neuron subset. We present NeuronGuard, a fine-tuning-stage defense that simultaneously hardens LLMs against both attack classes by redistributing safety signals across a broader set of neurons. NeuronGuard dynamically identifies safety-critical neurons via periodically refreshed per-layer linear classifiers, forces refusal behavior under deliberate neuron ablation, and applies KL-divergence regularization for distributional consistency. A randomized gradient projection strategy preserves downstream task utility by resolving conflicts between the defense and task objectives. We provide a formal guarantee that NeuronGuard strictly reduces the attack success rate (ASR) upper bound, and experiments across three LLMs, six state-of-the-art attack strategies, and multimodal settings confirm near-zero ASR while maintaining task accuracy, including against white-box adaptive adversaries.
CommentsTo appear in EMNLP 2026 (Findings)