arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14392cs.AI

Tripwire:通过统计验证的安全神经元触发对齐后的拒绝行为

Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons

Wei Zhao, Zhe Li, Peixin Zhang, Jun Sun

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出无需训练的Tripwire防御方法,通过统计验证安全神经元触发对齐拒绝,将攻击成功率降至最高2.0%,MT-Bench实用性下降仅0.5%-5.3%,为最优防御之一。

中文摘要 AI 辅助

神经元级和路径级干预是防御大型语言模型(LLM)抵御越狱攻击的最细粒度途径,但现有方法未能实现这一目标,往往会显著损害模型的实用性。具体而言,一类工作通过抑制毒性神经元来消除有害语义,但由于此类语义分布在整个网络中,阻断每条路径会导致干预范围过大;另一类研究则使用外部分类器识别安全神经元,虽有前景,但现有方法会损害对模型实用性重要的神经元。此外,两种方法均持续生效,即使无攻击存在也会干扰每个良性请求。为解决这些局限,本文提出Tripwire,一种无需训练的防御方法:首先通过每神经元假设检验(在错误发现率控制下)结合实用性特异性过滤器识别安全特定神经元;基于此,触发式钳位将选定神经元保持在有害条件下的平均激活值,注入内部有害输入信号以触发对齐过程中学习到的拒绝行为。该钳位通过两种可证等价的部署模式实现:检测器门控的推理时干预和离线偏差补丁权重编辑。在四个安全对齐LLM和四种代表性攻击上的大量实验表明,Tripwire将平均攻击成功率降至最高2.0%,同时在MT-Bench上仅产生0.5%至5.3%的实用性下降,是所有防御方法中最小的。代码可在该https URL获取。

英文摘要

Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this promise, i.e., they often compromise model utility significantly. Specifically, one line of work suppresses toxic neurons to erase harmful semantics, but since such semantics are distributed across the network, blocking every pathway forces a large intervention footprint. An alternative line of research focus on identify safety neurons using external classifiers. While promising, the existing approaches suffer from compromising neurons that are important for the model utility as well. Moreover, both approaches remain always on and thus perturb every benign request even when no attack is present. To address these limitations, we present \ours{}, a training-free defense that first identifies safety-specific neurons through per-neuron hypothesis tests under false-discovery-rate control together with a utility-specificity filter. Based on this identification, a trigger-style clamp holds the selected neurons at their harmful-conditional mean activations, injecting an internal harmful-input signal that triggers the refusal behavior learned during alignment. The clamp is then realized by two provably equivalent deployment modes, namely a detector-gated inference-time intervention and an offline bias-patch weight edit. Extensive experiments across four safety-aligned LLMs and four representative attacks demonstrate that \ours{} reduces the average attack success rate to at most 2.0\% while incurring a utility drop of only 0.5\% to 5.3\% on MT-Bench, the smallest among all defenses. Code is available at https://anonymous.4open.science/r/Tripwire-65C4.

↑