arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38140cs.CVcs.AI

打破均匀性陷阱:通过SplitMoE扩展视频扩散模型

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang, Yizhi Wang, Xinwei Huang, Minxuan Lin, Angtian Wang, Chongyang Ma, Fan Tang

首次发表
浏览论文内容

中文总结 AI 辅助

针对传统MoE在视频数据上的均匀性陷阱,提出SplitMoE,通过分裂专家池为语义与通用专家,结合原型引导路由和推拉正则化,在等效参数下提升收敛、路由连贯性和视频生成质量。

中文摘要 AI 辅助

混合专家(Mixture-of-Experts, MoE)范式因大型语言模型而流行,是扩展视觉生成模型的一种有前景的方法。然而,传统的词元级(token-wise)MoE在同质专家池中独立路由词元,并通过正则化使专家使用趋于均匀,这与在时空上冗余且语义上呈长尾分布的视频数据不匹配。我们表明,现有视觉MoE陷入了一个均匀性陷阱:语义组织不足的路由,加上均匀专家使用的正则化,将连贯的补丁分散到不同的专家中,导致路由碎片化和结构失真。为解决这一问题,我们提出SplitMoE,一种分裂角色的稀疏架构,打破了均匀性的束缚。为适应固有的语义不平衡,我们明确将专家池分为语义专家和通用专家,其中语义专家捕获高层语义抽象,通用专家保留残差视觉信息和灵活的生成能力。利用原型引导路由和推拉(pull-push)正则化,SplitMoE使词元能够根据语义属性自然聚类,而非受任意平衡约束。大量结果表明,在等效激活参数预算下,SplitMoE在标准基准上的收敛速度、路由连贯性和视频生成质量方面均优于传统的负载均衡MoE。通过揭示一种涌现的从粗到细的去噪逻辑,SplitMoE为社区提供了一条模态感知的扩展路径,为构建大规模视频世界模型提供了关键参考。

英文摘要

Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.

发表机构

  • University of Chinese Academy of Sciences(中国科学院大学)
  • ByteDance(字节跳动)
  • Canva Research(Canva 研究院)
  • University of Science and Technology Beijing(北京科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑