arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17402cs.CV

MoE-ViE:用于高效图像与视频理解的混合专家视觉编码器

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

Bonan Zhang, Shiyu Dong, Quan Hung Tran, Katharina Gschwind, Shuqi Yang, Sijia Chen, Adel Ahmadyan, Seungwhan Moon, Lu Zhang, Ahmed Kirmani, Babak Damavandi, Anuj Kumar

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出MoE-ViE,通过细粒度MoE拓扑、无辅助损失平衡变体、专用MoE内核及帧级蒸馏等设计,实现高效图像与视频理解,性能优于对应密集模型及更大规模SOTA编码器。

中文摘要 AI 辅助

视觉编码器是视觉-语言模型的关键组件,有效扩展其容量可提升性能,但密集扩展会增加计算成本与推理延迟。混合专家(MoE)架构在大语言模型(LLM)中已实现高效扩展,然而CLIP式视觉编码器的MoE设计空间在最先进(SOTA)水平下仍未得到充分探索。本研究系统探究用于视觉编码器扩展的MoE设计,发现细粒度MoE拓扑结构相比密集型与标准MoE均有显著增益;进一步提出无辅助损失的平衡变体以优化专家利用率,并设计专用MoE内核缓解推理延迟开销。为在保留图像知识的同时增强视频能力,引入帧级蒸馏与新型冻结机制。我们预训练了一系列不同规模的混合专家视觉编码器(MoE-ViE),均持续优于对应密集型模型;其中最大模型的零样本性能达到规模为其1.7倍的SOTA编码器水平,推理延迟仅为后者的76%。当与LLM对齐后,MoE-ViE在图像与视频基准测试中超越所有对比编码器,包括激活参数多至5倍的模型。代码可在指定URL获取。

英文摘要

Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and inference latency. Mixture-of-Experts (MoE) architectures offer a compelling alternative, having enabled efficient scaling in LLMs, yet the MoE design space for CLIP-style vision encoders remains underexplored at State-of-the-Art (SOTA) levels. In this work, we systematically study MoE designs for vision encoder scaling and find that fine-grained MoE topologies yield substantial gains over both dense and standard MoE counterparts. We further propose an auxiliary-loss-free balancing variant for better expert utilization, and design a specialized MoE kernel to mitigate inference latency overhead. To enhance video capabilities while preserving image knowledge, we introduce frame-level distillation paired with a novel freezing mechanism. We pretrain a series of Mixture-of-Experts Vision Encoders (MoE-ViE) across a range of sizes, all consistently outperforming their dense counterparts. Our largest model matches the zero-shot performance of a SOTA encoder 1.7x its size at 76% of its latency. When aligned with an LLM, MoE-ViE surpasses all compared encoders on image and video benchmarks, including those with up to 5x more activated parameters. Code is available at https://github.com/facebookresearch/moe_vie.

发表机构

  • Meta

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑