发表机构
Fudan University; The Chinese University of Hong Kong(复旦大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DecoMoE通过解耦视觉传播与专家计算,利用样本自适应视觉边界和路由校准专家前缀,在保持97.91%性能的同时实现1.69倍加速。
AI 中文摘要
多模态混合专家(MoE)模型将稀疏专家激活与视觉-语言能力相结合,但其推理成本仍然高昂,因为长视觉令牌序列会反复引发注意力、路由、分发和专家MLP计算。现有方法通常压缩令牌或专家维度中的某一个,而另一个维度仍存在冗余。我们的分析揭示了两种互补的规律:视觉传播所需的深度因输入而异,而文本令牌路由则表现出集中且循环的专家重要性模式。基于这些观察,我们提出DecoMoE,一种二维结构化压缩框架,将视觉传播与专家计算解耦。样本自适应视觉边界(SAVB)预测一个依赖于输入的视觉退出层,在该层移除视觉令牌块。路由校准专家前缀(RCEP)利用文本令牌路由质量离线重新排序专家,并从该预测的退出层起,在每个MoE层保留覆盖目标路由质量分数的最短连续前缀。我们在Qwen3-VL-MoE和InternVL3.5-30B-A3B上评估了DecoMoE,涵盖六个基准。在Qwen3-VL-MoE上,DecoMoE保留了密集基线性能的97.91%,同时将计算量从27.06 TFLOPs降至16.73 TFLOPs,延迟从0.44秒降至0.26秒,实现了1.69倍的加速。代码将在此https URL提供。
英文摘要
Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly because long visual-token sequences repeatedly incur attention, routing, dispatch, and expert-MLP computation. Existing methods typically compress either the token or expert dimension, leaving redundancy along the other. Our analysis reveals two complementary regularities: the depth required for visual propagation varies across inputs, while text-token routing exhibits concentrated and recurrent expert-importance patterns. Based on these observations, we propose DecoMoE, a two-dimensional structured compression framework that decouples visual propagation from expert computation. The Sample-Adaptive Visual Boundary (SAVB) predicts an input-dependent visual-exit layer at which the visual-token block is removed. The Routing-Calibrated Expert Prefix (RCEP) reorders experts offline using text-token routed mass and, from this predicted exit layer onward, retains at each MoE layer the shortest contiguous prefix covering a target routed-mass fraction. We evaluate DecoMoE on Qwen3-VL-MoE and InternVL3.5-30B-A3B across six benchmarks. On Qwen3-VL-MoE, DecoMoE retains 97.91% of dense-baseline performance while reducing computation from 27.06 to 16.73 TFLOPs and latency from 0.44 to 0.26 seconds, yielding a 1.69x speedup. Code will be available at https://github.com/ShawnTan86/DecoMoE.
Comments21 pages, 11 figures, 6 tables