MoRoute:面向上下文多模态视频生成的动态路由
MoRoute: Dynamic Routing for In-Context Multimodal Video Generation
- Sun Yat-sen University(中山大学)
- HUJING Digital Media & Entertainment Group(沪景数字媒体娱乐集团)
- Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MoRoute是采用动态层路由连接VLM与视频DiT的多模态视频生成框架,在三个基准数据集上均优于现有最优方法,提升了视频生成与编辑的性能。
AI中文摘要:
多模态视频生成旨在在单个模型中生成和编辑以文本、图像、视频任意组合为条件的视频,使不同任务能够共享互补数据与生成先验。统一这些任务需要对多样条件进行多模态理解,通常由预训练视觉语言模型(VLM)提供支持。核心挑战在于如何将VLM的分层多模态表示与预训练视频扩散Transformer(DiT)相连接。现有方法要么仅注入VLM最终层或少数手动选定层的特征,要么联合训练架构匹配的理解与生成流,难以复用异构预训练骨干网络。我们提出MoRoute,一种统一多模态视频生成框架,将冻结的VLM与架构不同的预训练视频DiT视为异构专家,通过动态层路由实现连接。对于每个输入,轻量的分块路由器使每个DiT块能选择与其生成阶段最相关的VLM层,从而学习多模态理解与视频合成间的自适应对应关系。MoRoute还通过统一上下文条件直接将参考图像与源视频融入DiT token序列,在各类生成与编辑任务中保留细粒度视觉细节。在IntelligentVBench、OpenVE-Bench和RefVIE-Bench上的实验表明,MoRoute在各基准上均优于最优对比方法,在1-5分制下平均得分分别提升0.15、0.18和0.34。
英文摘要:
Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, allowing diverse tasks to share complementary data and generative priors. Unifying these tasks requires multimodal understanding of diverse conditions, which is typically provided by a pretrained vision-language model (VLM). A key challenge is how to connect the VLM's hierarchical multimodal representations with a pretrained video diffusion transformer (DiT). Existing methods either inject features from only the final or a few manually selected VLM layers, or jointly train architecture-matched understanding and generation streams, making it difficult to reuse heterogeneous pretrained backbones. We introduce MoRoute, a unified multimodal video generation framework that formulates a frozen VLM and a pretrained video DiT with different architectures as heterogeneous experts connected through dynamic layer routing. For each input, a lightweight block-wise router enables every DiT block to select the VLM layer most relevant to its generation stage, thereby learning an adaptive correspondence between multimodal understanding and video synthesis. MoRoute further incorporates reference images and source videos directly into the DiT token sequence through unified in-context conditioning, preserving fine-grained visual details across diverse generation and editing tasks. Experiments on IntelligentVBench, OpenVE-Bench, and RefVIE-Bench show that MoRoute consistently surpasses the best competing method on each benchmark, improving the average score by 0.15, 0.18, and 0.34 on a 1-5 scale, respectively.