arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mask2Shield:增强大语言模型抵御神经元剪枝攻击的安全性

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks

Ying JinCheng, Minghui Xu, Yinhao Xiao, Xiuzhen Cheng, Wencheng Yang

arXiv 2607.23015首次发表:更新:

AI 中文总结

研究针对大语言模型神经元剪枝攻击问题,提出Mask2Shield方法,通过掩码前向对齐训练模型,减少对可移除安全神经元集的依赖,降低针对性剪枝攻击成功率,同时保持能力基准。

AI 中文摘要

大语言模型在部署前进行安全对齐以减少有害内容生成。然而,神经元级剪枝攻击表明,模型的拒绝行为可能依赖于一小部分可移除的单元:禁用它们会消除安全行为,而模型的大部分仍可使用。为解决此问题,我们引入了Mask2Shield(M2S),一种在功能剪枝下训练模型的掩码前向对齐方法。掩码后的学生模型必须通过剩余计算恢复安全拒绝,而冻结的未掩码教师模型提供完整的良性答案以限制能力漂移。在十种模型配置中,M2S将313个提示中成功的重新计算剪枝攻击从80 - 279次减少到1 - 44次,同时总体上保持了四个能力基准。我们还使用TwinBreak对M2S进行了评估,TwinBreak使用不同的神经元选择规则和迭代剪枝过程。这些结果共同表明,M2S通过减少对一小部分可移除安全神经元集的依赖,使针对性剪枝的效果降低。

英文摘要

Large language models (LLMs) are safety-aligned before deployment to reduce harmful content generation. Yet neuron-level pruning attacks show that refusal can depend on a small set of removable units: disabling them can remove safety behavior while leaving much of the model usable. To address this problem, we introduce Mask2Shield (M2S), a masked-forward alignment method that trains a model under this functional pruning. The masked student must recover a safe refusal through the remaining computation, while a frozen, unmasked teacher supplies complete benign answers to limit capability drift. Across ten model configurations, M2S reduces successful recomputed pruning attacks from 80--279 to 1--44 out of 313 prompts while generally preserving four capability benchmarks. We also evaluate M2S with TwinBreak, which uses a different neuron-selection rule and iterative pruning procedure. Together, these results show that M2S makes targeted pruning less effective by reducing reliance on a small, removable safety-neuron set.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑