arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多网格后训练用于长格式多镜头视频生成

Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

Jiawei Mao, Haoqin Tu, Hardy Chen, Yuhan Wang, Keyang Xu, Jieru Mei, Hongliang Fei, Ruogu Fang, Wei Shao, Cihang Xie, Yuyin Zhou

arXiv 2609.06373首次发表:更新:

发表机构

UC Santa Cruz; Google; University of Florida; Vanderbilt University(加州大学圣克鲁兹分校; 谷歌; 佛罗里达大学; 范德堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MovieGrid通过多网格后训练将长视频分解为空间网格块,实现高效长格式多镜头生成,在一致性和镜头数量上显著超越现有方法。

AI 中文摘要

生成长格式多镜头视频需要镜头内运动连贯以及跨镜头的视觉一致叙事。现有视频生成器倾向于连续运动,当整个叙事沿单一时间轴打包时,难以呈现完整的镜头集合。我们提出MovieGrid,一种多网格后训练范式,将长视频分解为较短、按时间排序的块,并将它们排列在空间网格上进行联合建模。这种设计减少了每个时间轴处理的镜头数量,同时实现跨块全局信息交换。我们利用源视频收集、层次分割、网格视频构建和角色感知故事标注,从1,000个长视频构建了多网格长视频(MGLV)数据集,生成了54K个带故事提示的网格视频。我们的无噪声随机网格训练保留随机子集的块作为干净视觉上下文,用于去噪其余块。网格嵌入编码网格结构,角色感知故事提示关联重复出现的实体,网格边界损失稳定布局。在相同token预算下,MovieGrid在1,616帧视频中生成的镜头数量是时间打包的6.05倍。在涵盖五个真实世界类别的基准上,它实现了镜头内一致性(0.9131对比HoloCine的0.8086)和镜头间一致性(0.5914对比StoryMem的0.5384)的最先进性能。MovieGrid可以通过单次或多次生成进一步扩展视频长度,且影响极小。

英文摘要

Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across shots. Existing video generators favor continuous motion and struggle to present complete shot sets when an entire narrative is packed along one temporal axis. We propose MovieGrid, a Multi-Grid Post-Training paradigm that decomposes a long video into shorter, temporally ordered chunks and arranges them on a spatial grid for joint modeling. This design reduces the number of shots handled by each temporal axis while enabling global information exchange across chunks. We construct the Multi-Grid Long Video (MGLV) dataset from 1,000 long-form videos using source video collection, hierarchical segmentation, grid video construction, and character-aware story annotation, producing 54K grid videos paired with story prompts. Our Noise-Free Random-Grid Training retains a random subset of chunks as clean visual context for denoising the remaining chunks. Grid Embedding encodes grid structure, character-aware Story Prompts link recurring entities, and Grid Boundary Loss stabilizes layouts. Under the same token budget, MovieGrid generates 6.05 times more shots than Temporal Packing in a 1,616-frame video. On a benchmark spanning five real-world categories, it achieves state-of-the-art intra-shot consistency (0.9131 versus 0.8086 for HoloCine) and inter-shot consistency (0.5914 versus 0.5384 for StoryMem). MovieGrid can further scale video length with minimal compromise through single or multiple generations.

Comments17 pages, 13 figures. Project page: https://jwmao1.github.io/moviegrid_web

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑