arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DecoMoE:解耦视觉传播与专家计算以实现高效多模态MoE推理

DecoMoE: Decoupling Visual Propagation and Expert Computation for Efficient Multimodal MoE Inference

Xudong Tan, Peng Ye, Ming Xie, Chenyu Huang, Yaoxin Yang, Jiayuan Fan, Tao Chen

arXiv 2609.38823首次发表:更新:

发表机构

Fudan University; The Chinese University of Hong Kong(复旦大学; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DecoMoE通过解耦视觉传播与专家计算,利用样本自适应视觉边界和路由校准专家前缀,在保持97.91%性能的同时实现1.69倍加速。

AI 中文摘要

多模态混合专家(MoE)模型将稀疏专家激活与视觉-语言能力相结合,但其推理成本仍然高昂,因为长视觉令牌序列会反复引发注意力、路由、分发和专家MLP计算。现有方法通常压缩令牌或专家维度中的某一个,而另一个维度仍存在冗余。我们的分析揭示了两种互补的规律:视觉传播所需的深度因输入而异,而文本令牌路由则表现出集中且循环的专家重要性模式。基于这些观察,我们提出DecoMoE,一种二维结构化压缩框架,将视觉传播与专家计算解耦。样本自适应视觉边界(SAVB)预测一个依赖于输入的视觉退出层,在该层移除视觉令牌块。路由校准专家前缀(RCEP)利用文本令牌路由质量离线重新排序专家,并从该预测的退出层起,在每个MoE层保留覆盖目标路由质量分数的最短连续前缀。我们在Qwen3-VL-MoE和InternVL3.5-30B-A3B上评估了DecoMoE,涵盖六个基准。在Qwen3-VL-MoE上,DecoMoE保留了密集基线性能的97.91%,同时将计算量从27.06 TFLOPs降至16.73 TFLOPs,延迟从0.44秒降至0.26秒,实现了1.69倍的加速。代码将在此https URL提供。

英文摘要

Multimodal mixture-of-experts (MoE) models combine sparse expert activation with visual-language capabilities, yet their inference remains costly because long visual-token sequences repeatedly incur attention, routing, dispatch, and expert-MLP computation. Existing methods typically compress either the token or expert dimension, leaving redundancy along the other. Our analysis reveals two complementary regularities: the depth required for visual propagation varies across inputs, while text-token routing exhibits concentrated and recurrent expert-importance patterns. Based on these observations, we propose DecoMoE, a two-dimensional structured compression framework that decouples visual propagation from expert computation. The Sample-Adaptive Visual Boundary (SAVB) predicts an input-dependent visual-exit layer at which the visual-token block is removed. The Routing-Calibrated Expert Prefix (RCEP) reorders experts offline using text-token routed mass and, from this predicted exit layer onward, retains at each MoE layer the shortest contiguous prefix covering a target routed-mass fraction. We evaluate DecoMoE on Qwen3-VL-MoE and InternVL3.5-30B-A3B across six benchmarks. On Qwen3-VL-MoE, DecoMoE retains 97.91% of dense-baseline performance while reducing computation from 27.06 to 16.73 TFLOPs and latency from 0.44 to 0.26 seconds, yielding a 1.69x speedup. Code will be available at https://github.com/ShawnTan86/DecoMoE.

Comments21 pages, 11 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑