arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越全局路由聚合:面向MoE视觉-语言模型的阶段感知专家合并

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models

Hongyu Zhang, Cheng Yan, Xiang Xia, Wuyang Zhang

arXiv 2608.04454首次发表:更新:

AI 中文总结

针对MoE-VLM全局路由聚合导致专家角色混淆的问题,提出无需训练的RoleMerge方法,通过阶段归一化路由统计合并专家,在匹配专家保留比例下性能优于现有方法,六任务宏平均相对提升最高9.6%。

AI 中文摘要

混合专家视觉-语言模型(MoE-VLMs)通过稀疏专家激活提升模型容量,但部署需存储完整专家池。无需训练的专家合并可减轻该负担,许多基于路由的方法会聚合所有token的路由统计以确定合并兼容性。然而,MoE-VLM推理具有阶段结构:图像上下文token承载视觉内容,问题token指定查询,答案token生成输出,三者数量和路由分布不同。由于图像上下文token数量多得多,全局聚合会过度强调图像上下文处理,掩盖阶段相关的专家角色,导致服务不同阶段的专家看起来可互换,降低模型性能。因此,我们认为MoE-VLM的专家合并应保留阶段相关的专家角色,通过专家服务不同阶段的方式而非全局聚合的路由统计来判断兼容性。基于此观点,我们提出RoleMerge,一种无需训练的方法,它从阶段归一化的路由统计中构建每个专家的路由角色轮廓(RRP),捕捉其相对阶段偏好。在专家-阶段信息损失的引导下,RoleMerge合并具有兼容轮廓的专家及其对应的路由条目,同时保留答案解码专家的区分度。在三个模型和多个基准上的实验表明,在匹配的专家保留比例下,RoleMerge比其他专家合并方法保留了更多完整模型的性能,在六任务宏平均性能上实现了高达9.6%的相对提升。这些结果验证了阶段相关的专家角色是比全局路由聚合更有效的MoE-VLM专家合并基础。

英文摘要

Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full expert pool. Training-free expert merging reduces this burden, and many routing-based methods aggregate routing statistics across all tokens to determine merge compatibility. However, MoE-VLM inference is phase-structured: image-context tokens carry visual content, question tokens specify the query, and answer tokens produce the output, with different counts and routing distributions. Because image-context tokens are far more numerous, global aggregation can overemphasize image-context processing and obscure phase-conditioned expert roles, making experts serving different phases appear interchangeable and degrading model performance. We therefore argue that MoE-VLM expert merging should preserve phase-conditioned expert roles, judging compatibility by how experts serve different phases rather than globally aggregated routing statistics. Based on this view, we propose RoleMerge, a training-free method that constructs each expert's Routing Role Profile (RRP) from phase-normalized routing statistics, capturing its relative phase preference. Guided by expert-phase information loss, RoleMerge merges experts with compatible profiles and their corresponding router entries while preserving answer-decoding expert distinctions. Experiments on three models and multiple benchmarks show that RoleMerge preserves more of the full model's performance than alternative expert-merging methods at matched expert-retention ratios, with relative improvements of up to 9.6 percent in six-task macro-average performance. These results validate phase-conditioned expert roles as a more effective basis than global routing aggregation for MoE-VLM expert merging.

Comments17 pages, 3 figures, 17 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑