MoE-ViE:用于高效图像与视频理解的混合专家视觉编码器
MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
浏览论文内容
中文总结 AI 辅助
本研究提出MoE-ViE,通过细粒度MoE拓扑、无辅助损失平衡变体、专用MoE内核及帧级蒸馏等设计,实现高效图像与视频理解,性能优于对应密集模型及更大规模SOTA编码器。
中文摘要 AI 辅助
视觉编码器是视觉-语言模型的关键组件,有效扩展其容量可提升性能,但密集扩展会增加计算成本与推理延迟。混合专家(MoE)架构在大语言模型(LLM)中已实现高效扩展,然而CLIP式视觉编码器的MoE设计空间在最先进(SOTA)水平下仍未得到充分探索。本研究系统探究用于视觉编码器扩展的MoE设计,发现细粒度MoE拓扑结构相比密集型与标准MoE均有显著增益;进一步提出无辅助损失的平衡变体以优化专家利用率,并设计专用MoE内核缓解推理延迟开销。为在保留图像知识的同时增强视频能力,引入帧级蒸馏与新型冻结机制。我们预训练了一系列不同规模的混合专家视觉编码器(MoE-ViE),均持续优于对应密集型模型;其中最大模型的零样本性能达到规模为其1.7倍的SOTA编码器水平,推理延迟仅为后者的76%。当与LLM对齐后,MoE-ViE在图像与视频基准测试中超越所有对比编码器,包括激活参数多至5倍的模型。代码可在指定URL获取。
英文摘要
Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and inference latency. Mixture-of-Experts (MoE) architectures offer a compelling alternative, having enabled efficient scaling in LLMs, yet the MoE design space for CLIP-style vision encoders remains underexplored at State-of-the-Art (SOTA) levels. In this work, we systematically study MoE designs for vision encoder scaling and find that fine-grained MoE topologies yield substantial gains over both dense and standard MoE counterparts. We further propose an auxiliary-loss-free balancing variant for better expert utilization, and design a specialized MoE kernel to mitigate inference latency overhead. To enhance video capabilities while preserving image knowledge, we introduce frame-level distillation paired with a novel freezing mechanism. We pretrain a series of Mixture-of-Experts Vision Encoders (MoE-ViE) across a range of sizes, all consistently outperforming their dense counterparts. Our largest model matches the zero-shot performance of a SOTA encoder 1.7x its size at 76% of its latency. When aligned with an LLM, MoE-ViE surpasses all compared encoders on image and video benchmarks, including those with up to 5x more activated parameters. Code is available at https://github.com/facebookresearch/moe_vie.
发表机构
- Meta
机构由 AI 辅助整理,请以论文原文为准。