arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26043cs.LG

鲁棒CurveMoE:通过模式连通性实现混合专家模型的多范数对抗防御

Robust CurveMoE: Multi-Norm Adversarial Defense for Mixture-of-Experts Models via Mode Connectivity

Xu Zhang, Ren Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出Robust CurveMoE混合专家框架,通过模式连通性连接不同范数的专用模型,引入贡献引导的部分更新降低成本,在CIFAR-100等数据集上显著提升了各类准确率。

中文摘要 AI 辅助

多范数对抗防御旨在保护神经网络免受不同范数约束定义的扰动影响,但现有方法通常在单一参数配置下优化相互竞争的鲁棒性目标,导致训练成本高昂且鲁棒性权衡效果不佳。我们提出Robust CurveMoE,这是一种高效的混合专家(Mixture-of-Experts)框架,通过低损失路径连接针对不同扰动范数的专用模型,并利用该路径上模型的互补鲁棒性特征。Robust CurveMoE从受鲁棒性约束的曲线位置提取干净的、范数专用的专家,仅选择性地对有影响力的层进行专家化,同时在路由路径间共享其余参数。为进一步降低曲线构建成本,我们引入贡献引导的部分更新方法,该方法利用基于初始化的梯度分数选择有影响力的曲线参数。我们还从理论上界定了部分曲线优化与全曲线优化之间的目标差距。在CIFAR-100和ImageNet-100数据集上,使用WideResNet和Vision Transformer架构进行的实验表明,Robust CurveMoE在干净准确率、范数专用准确率和联合准确率(Union accuracy)上均优于MSD和ERMC。尤其在CIFAR-100和ImageNet-100上,它分别比最强基线方法将联合准确率提升了2.37和2.13个百分点。大量 ablation 实验进一步验证了部分更新、选择性专家化以及受鲁棒性约束的专家选择的有效性。

英文摘要

Multi-norm adversarial defense aims to protect neural networks against perturbations defined by different norm constraints, but existing methods typically optimize competing robustness objectives within a single parameter configuration, leading to substantial training cost and unfavorable robustness trade-offs. We propose Robust CurveMoE, an efficient mixture-of-experts framework that connects models specialized for different perturbation norms through a low-loss path and exploits the complementary robustness profiles of models along this path. Robust CurveMoE derives clean and norm-specialized experts from robustness-constrained curve locations and selectively expertizes only influential layers, while sharing the remaining parameters across routing paths. To further reduce curve-construction cost, we introduce contribution-guided partial updating, which selects influential curve parameters using initialization-based gradient scores. We also theoretically bound the objective gap between partial and full curve optimization. Experiments on CIFAR-100 and ImageNet-100 with WideResNet and Vision Transformer architectures show that Robust CurveMoE consistently improves clean, norm-specific, and Union accuracy over MSD and ERMC. In particular, it improves Union accuracy by 2.37 and 2.13 percentage points over the strongest baseline on CIFAR-100 and ImageNet-100, respectively. Extensive ablations further validate the effectiveness of partial updating, selective expertization, and robustness-constrained expert selection.

发表机构

  • Illinois Institute of Technology(伊利诺伊理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑