arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

频率不等于敏感性:识别稀疏MoE大语言模型中的安全敏感专家

Frequency Is Not Sensitivity Identifying Safety-Sensitive Experts in Sparse MoE LLM

Md Nurul Absar Siddiky, Liuwan Zhu, Yingfei Dong

arXiv 2610.02910首次发表:更新:

发表机构

University of Hawaii at Manoa(夏威夷大学马诺阿分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出用路由器梯度敏感性替代激活频率来识别稀疏MoE大语言模型中的安全敏感专家,实验表明该方法在多种架构和预算下能更有效地降低恶意提示的拒绝率,且不损害生成质量。

AI 中文摘要

在不进行重新训练的情况下,抑制一小部分路由专家可以削弱稀疏混合专家(MoE)语言模型的安全行为。因此,抑制哪些专家是一个安全问题,通常的答案是激活频率,但频率衡量的是使用情况,而非影响力。我们测试了一种替代方案:路由器梯度敏感性,即序列损失对选择专家的门控权重的敏感性。在五种MoE架构中,我们根据500个良性提示和500个恶意提示对每个信号进行专家排名,并在两种预算下测量100个保留恶意提示的拒绝率:专家数量相等和名义恶意路由流量相等(1%-5%)。在两种预算下,路由器梯度选择在25种条件中的24种中比激活更能降低拒绝率,并且在所有25种条件中比十次随机试验的平均值更能降低拒绝率。最大的效果出现在OLMoE中,拒绝率从100个提示中的34个降至9个(相对降低73.53%),且没有输出质量下降,表明是实质性的顺从而非生成损坏。在匹配每一层的专家数量后,梯度选择在25种条件中的23种中仍比激活产生更大的拒绝率降低,另有2种条件持平。一项探索性的跨模型分析将更大的恶意与良性集中度差距与更大的峰值梯度效应联系起来(rho = 0.90;精确双侧p = 0.083,n = 5)。总之,结果支持在所测试的预算下使用梯度选择。

英文摘要

Suppressing a small set of routed experts can weaken the safety behavior of a sparse Mixture-of-Experts (MoE) language model without retraining. Which experts to suppress is therefore a security question, and the usual answer is activation frequency, but frequency measures use, not influence. We test an alternative: router-gradient sensitivity, the sensitivity of the sequence loss to the gate weights that select an expert. Across five MoE architectures, we rank experts by each signal on 500 benign and 500 malicious prompts and measure refusal on 100 held-out malicious prompts under two budgets: equal expert counts and equal nominal malicious routing traffic (1%-5%). Under each of the two budgets, router-gradient selection reduces refusals more than activation in 24 of 25 conditions, and more than a ten-trial random mean in all 25. The largest effect is in OLMoE, where refusals fall from 34 to 9 of 100 prompts (73.53% relative) with no degraded outputs, indicating substantive compliance rather than broken generation. After matching expert counts in every layer, gradient selection still produces greater refusal reduction than activation in 23 of 25 conditions, with two ties. An exploratory cross-model analysis links larger malicious-versus-benign concentration gaps to greater peak gradient effects (rho = 0.90; exact two-sided p = 0.083, n = 5). Together, the results support gradient selection under the tested budgets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑