发表机构
The University of Sydney; National University of Singapore(悉尼大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有安全对齐方法的单点失效问题,提出分布式安全对齐(DSA),通过冗余编码安全能力提升模型对抗白盒神经元级攻击的鲁棒性,且保留通用语言与多模态效用。
AI 中文摘要
随着开源权重大型基础模型的快速发布,安全威胁正从黑盒越狱转向直接识别和操纵安全相关神经元的神经元级白盒攻击。现有对齐方法通常仅在少量神经元上研究安全行为,形成冗余有限的脆弱单点失效。为解决该问题,我们提出分布式安全对齐(DSA),其将安全能力冗余编码到多个计算神经元中,确保即使关键安全神经元被破坏,模型仍能维持安全基线。具体而言,我们将干预定位于语言侧前馈网络中下投影层的输入,并将每个特征坐标视为单个神经元的激活。随后,DSA结合神经元激活与损失梯度,计算方向感知的一阶泰勒分数,以全局识别对模型当前弃权(不执行)行为贡献最大的神经元。最后,通过确定性掩码和随机丢弃进行针对性破坏,迫使模型放弃狭窄的安全神经元,并将安全行为冗余编码到多个补偿神经元中。大量实验表明,DSA在保持模型通用语言和多模态效用的同时,显著提升了对抗白盒神经元级安全攻击的鲁棒性。
英文摘要
With the rapid release of open-weight large foundation models, safety threats are shifting from black-box jailbreaks to neuron-level white-box attacks that directly identify and manipulate safety-related neurons. Existing alignment methods often investigate the safety behavior on a small number of neurons, creating fragile single point of failure with limited redundancy. To address this issue, we propose distributed safety alignment (DSA), which redundantly encodes safety capabilities across multiple computational neurons, ensuring that the model maintains its safety baseline even when critical safety neurons are disrupted. Specifically, we localize the intervention to the inputs of the down-projection layers in language-side feed-forward networks and treat each feature coordinate as the activation of an individual neuron. DSA then combines neuron activations with loss gradients to compute a direction-aware first-order Taylor score that globally identifies the neurons that contribute most to the current refusal behavior of the model. Finally, targeted disruption via deterministic masking and stochastic dropout is coupled, forcing the model to abandon narrow safety neurons and redundantly encode safety behavior across multiple compensatory neurons. Extensive experiments show that DSA substantially improves robustness against white-box neuron-level safety attacks while preserving the model's general language and multimodal utility.