发表机构
Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对混合专家(MoE)模型易受对抗攻击的问题,提出基于共享专家对齐的参数高效防御方法SEAL及变体SEAL++,可降低攻击成功率且性能损失极小。
AI 中文摘要
混合专家(MoE)是大语言模型的扩展架构,每个token仅激活小部分专家模块,能在计算量几乎恒定的情况下实现参数规模的大幅增长。近期的混合MoE架构新增了共享专家以捕获持续有用的表征,进一步提升了稳定性与泛化能力。MoE现已驱动众多旗舰级开源及商业模型,但仍易受对抗攻击。具体而言,稀疏路由会引入结构性漏洞:MoE的安全性取决于激活哪些专家,攻击者可通过越狱提示、恶意微调、对安全关键神经元的权重级剪枝来破坏这种选择。现有防御措施主要聚焦于强化路由机制,但由于路由过程的非确定性,攻击者仍可能操纵或绕过路由轨迹,导致防御失效。为解决该问题,我们从理论和实证层面首次发现:共享专家(一种始终激活的组件,包含小比例安全关键神经元)可克服稀疏激活路由路径的不确定性,作为与路由无关的锚点来增强全局安全对齐。基于此见解,我们提出SEAL,一种训练时的参数高效防御方法,生成一个可即插即用的适配器附加到共享专家;还提出其变体SEAL++,在训练过程中添加正交约束以保留预存在的安全子空间。我们在六种攻击场景中评估了SEAL和SEAL++,这些场景结合了三种对抗输入(有害提示、越狱、恶意微调)以及有无神经元剪枝的情况。SEAL可将攻击成功率(ASR)降低最多60%,在五个基准的平均性能上最多仅产生1.4%的能力损失。此外,SEAL可与路由层无缝集成。
英文摘要
Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE architecture adds \textit{shared experts} to capture consistently useful representations, further improving stability and generalization. MoE now powers many flagship open-source and commercial models, yet remains vulnerable to adversarial attacks. Specifically, sparse routing introduces a structural vulnerability: MoE safety hinges on which experts are activated, and adversaries can subvert this selection through jailbreak prompts, malicious fine-tuning, and weight-level pruning of safety-critical neurons. Existing defenses primarily focus on hardening the router, but an adversary may still manipulate or bypass the routing trajectory due to the routing process's nondeterministic nature, thereby collapsing the defense. To cope with this problem, we first identify theoretically and empirically that shared expert, an always-activated component containing a small proportion of safety-critical neurons, can overcome the uncertainty of sparsely activated routing path and serve as a router-independent anchor to enhance global safety alignment. Based on this insight, we propose SEAL, a training-time parameter-efficient defense that produces a plug-and-play adapter attached to shared expert, and SEAL++, a variant that adds an orthogonal constraint preserving pre-existing safety subspaces during training. We evaluate SEAL and SEAL++ across six attack scenarios that combine three adversarial inputs (harmful prompting, jailbreak, malicious fine-tuning) with and without neuron pruning. SEAL reduces attack success rate (ASR) by up to 60\%, at a capability cost of at most 1.4\% on a five-benchmark average. Additionally, SEAL can seamlessly integrate with router-level ......
CommentsAccepted at ACM CCS 2026