AI 中文总结
研究针对大语言模型神经元剪枝攻击问题,提出Mask2Shield方法,通过掩码前向对齐训练模型,减少对可移除安全神经元集的依赖,降低针对性剪枝攻击成功率,同时保持能力基准。
AI 中文摘要
大语言模型在部署前进行安全对齐以减少有害内容生成。然而,神经元级剪枝攻击表明,模型的拒绝行为可能依赖于一小部分可移除的单元:禁用它们会消除安全行为,而模型的大部分仍可使用。为解决此问题,我们引入了Mask2Shield(M2S),一种在功能剪枝下训练模型的掩码前向对齐方法。掩码后的学生模型必须通过剩余计算恢复安全拒绝,而冻结的未掩码教师模型提供完整的良性答案以限制能力漂移。在十种模型配置中,M2S将313个提示中成功的重新计算剪枝攻击从80 - 279次减少到1 - 44次,同时总体上保持了四个能力基准。我们还使用TwinBreak对M2S进行了评估,TwinBreak使用不同的神经元选择规则和迭代剪枝过程。这些结果共同表明,M2S通过减少对一小部分可移除安全神经元集的依赖,使针对性剪枝的效果降低。
英文摘要
Large language models (LLMs) are safety-aligned before deployment to reduce harmful content generation. Yet neuron-level pruning attacks show that refusal can depend on a small set of removable units: disabling them can remove safety behavior while leaving much of the model usable. To address this problem, we introduce Mask2Shield (M2S), a masked-forward alignment method that trains a model under this functional pruning. The masked student must recover a safe refusal through the remaining computation, while a frozen, unmasked teacher supplies complete benign answers to limit capability drift. Across ten model configurations, M2S reduces successful recomputed pruning attacks from 80--279 to 1--44 out of 313 prompts while generally preserving four capability benchmarks. We also evaluate M2S with TwinBreak, which uses a different neuron-selection rule and iterative pruning procedure. Together, these results show that M2S makes targeted pruning less effective by reducing reliance on a small, removable safety-neuron set.