arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

前沿规模下安全对齐有多脆弱?针对320B MoE的单方向攻击

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

Yi Shi, Tanyu Chen, Kai Shen

arXiv 2609.09793首次发表:更新:

发表机构

Continuum AI(Continuum人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究前沿规模MoE模型(320B GLM-5.3-Flash)上方向消融攻击的有效性,发现其效果主要依赖联合干预,且存在类别集中残留,在七个有害基准上实现41-89个百分点降幅。

AI 中文摘要

方向消融通过从写入残差流的权重中投影出单一的“拒绝方向”,从而移除对齐语言模型的拒绝能力。它不需要基于梯度的训练,也不需要优化,仅需几百个对比提示,这使其成为针对开放权重对齐的典型白盒攻击。然而,该方法仅在参数规模约70B的稠密模型上得到验证。我们研究该方法能否在残差流不再是单一张量且权重以量化形式发布的前沿混合专家(MoE)模型中继续有效。我们将其应用于GLM-5.3-Flash(320B参数,288个路由专家,四路超连接残差,块级FP8)。该攻击在架构上依然有效,但其作用位置已不再是原始方法读者所预期之处。单独编辑注意力、稠密和路由专家写入器分别移除0.039、0.016和0.148的拒绝能力;三者联合编辑则移除0.776。因此,74%的效果仅存在于联合干预之下。传统方法通过模块名匹配所触及的部分仅占0.776中的0.066,这正是其在MoE上静默失败的原因。该效果并非源于移除任意方向:消融与其正交的随机方向后,拒绝能力保持不变。一个类别集中的残留在所有编辑尝试后依然存在:针对暴力、色情内容和仇恨拟合的子空间在秩1至12的每个秩上均留下可测量的拒绝能力。我们报告了该方法、其在七个有害基准上实现的41至89个百分点的降幅(且未检测到能力变化),以及其失效的边界。

英文摘要

Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It needs no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-weight alignment. However, it has been established only on dense models up to roughly 70B parameters. We study whether it survives the shift to frontier mixture-of-experts (MoE) models whose residual streams are no longer a single tensor and whose weights ship quantized. We apply it to GLM-5.3-Flash (320B parameters, 288 routed experts, a four-wide hyper-connection residual, block-FP8). The attack survives the architecture, but what it reaches is no longer where a reader of the original recipe would look for it. Editing the attention, dense and routed-expert writers on their own removes 0.039, 0.016 and 0.148 of refusal respectively; editing all three together removes 0.776. As a result, 74% of the effect exists only under the joint intervention. The part the conventional recipe reaches by module-name matching accounts for 0.066 of that 0.776, which is why it fails silently on an MoE. The effect does not follow from removing just any direction: ablating a random direction orthogonal to it leaves refusal unchanged. A category-concentrated residue survives every edit we tried: subspaces fitted on violence, sexual content and hate leave measurable refusal at every rank from 1 to 12. We report the method, the 41-89 percentage-point reductions it achieves across seven harmful benchmarks with no detected change in capability, and the boundary where it stops.

Comments20 pages, 14 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑