超越文本条件:面向视频生成的MLLM-DiT融合系统研究
Beyond Text Conditioning: A Systematic Study of MLLM-DiT Fusion for Video Generation
浏览论文内容
中文总结 AI 辅助
该研究针对视频生成中MLLM与DiT融合问题,提出BiVidGen框架,通过MLLM生成语义视觉标记辅助DiT渲染,提升了视频的语义对齐度与时间一致性。
中文摘要 AI 辅助
扩散Transformer(DiT)已成为高保真视频生成的主流范式,但其高级语义规划能力仍有限。尽管将多模态大语言模型(MLLM)与扩散主干网络结合的混合架构在图像合成中展现出显著优势,但这类设计在视频生成领域的探索仍不充分,现有方法通常仅将MLLM作为冻结的特征编码器,而非语义生成器。为填补这一空白,本文通过回答三个问题系统研究MLLM应如何与DiT融合以用于视频生成:MLLM与DiT之间应采用何种中间表示、MLLM应如何生成该表示、DiT在扩散渲染过程中应如何整合该表示。分析得出三项关键发现:(1)基于EMA的分词器生成的离散语义视觉标记提供了稳定且具表达力的接口;(2)自回归因果建模对生成这些标记有效;(3)显式视觉标记条件比提示词优化或潜在空间桥接更有效。基于这些发现,本文提出BiVidGen框架,该混合框架中MLLM先生成语义视觉标记,DiT通过多层交叉注意力同时基于文本和这些标记作为条件渲染视频。大量实验表明,BiVidGen相比微调后的DiT基线提升了语义对齐度与时间一致性,在VBench-Long基准上取得了更优性能。这些结果证明,基于MLLM的显式视觉规划为文本到视频生成提供了超越纯文本条件的有效中间接口。
英文摘要
Diffusion Transformers (DiTs) have become the dominant paradigm for high-fidelity video generation, yet their ability to perform high-level semantic planning remains limited. While hybrid architectures integrating MLLMs with diffusion backbones have shown strong advantages in image synthesis, such designs remain underexplored in video generation, where existing approaches often treat MLLMs primarily as frozen feature encoders rather than semantic generators. To fill this gap, we systematically study how an MLLM should be integrated with a DiT for video generation by answering three questions: what intermediate representation should bridge the MLLM and DiT, how the MLLM should generate it, and how the DiT should incorporate it during diffusion rendering. Our analysis reveals three key findings: (1) discrete semantic visual tokens produced by an EMA-based tokenizer provide a stable and expressive interface, (2) autoregressive causal modeling is effective for generating these tokens, and (3) explicit visual-token conditioning is more effective than prompt refinement or latent bridging. Based on these findings, we propose BiVidGen, a hybrid framework where an MLLM first generates semantic visual tokens and a DiT renders videos conditioned on both text and these tokens via multi-layer cross-attention. Extensive experiments show that BiVidGen improves semantic alignment and temporal coherence over a fine-tuned DiT baseline, achieving stronger performance on VBench-Long. These results demonstrate that explicit MLLM-based visual planning provides an effective intermediate interface for text-to-video generation beyond text-only conditioning.
发表机构
- Chinese Academy of Sciences(中国科学院)
- Microsoft Research(微软研究院)
- Sun Yat-sen University(中山大学)
- Zhejiang University(浙江大学)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Xi’an Jiaotong University(西安交通大学)
机构由 AI 辅助整理,请以论文原文为准。