发表机构
Shanghai Jiao Tong University; China Telecom; Institute of Artificial Intelligence (TeleAI)(上海交通大学; 中国电信; 人工智能研究院(电信AI))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CineWeaver是首个无训练的统一框架,通过操控位置编码、注意力模式等,实现参考可控的多镜头长视频生成,产出的电影视频时长较长且质量高。
AI 中文摘要
电影视频生成对文本到视频扩散模型而言颇具挑战,因为需同时满足多镜头生成、角色与场景的细粒度可控性,以及跨长时序的长视频生成需求。现有方法依赖定制化与重训来分别应对特定需求,无法用统一框架同时满足所有要求。本文揭示无训练范式的关键洞见:多镜头生成的难点源于预训练视频扩散模型对时序连续性的结构偏差,进而提出名为CineWeaver的统一框架,实现无需重训的参考可控多镜头长视频生成。我们在推理阶段操控位置编码与注意力模式,打破时序连续性,使预训练视频扩散模型能生成清晰的镜头过渡;还为该框架扩展了镜头路由参考条件机制,实现单镜头细粒度可控性,并开发锚点记忆机制,支持带有一致全局外观线索的长视频生成。据我们所知,CineWeaver是首个以无训练方式同时实现长视频、参考可控与多镜头视频生成的统一框架。实验结果表明,CineWeaver可生成长时长、身份一致、全局外观稳定且镜头过渡清晰的高质量电影视频。项目页面可访问:this https URL
英文摘要
Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended temporal horizons. Existing methods rely on customization and retraining to separately address specific requirements, and cannot simultaneously fulfill all the requirements with a unified framework. In this paper, we shed light on the training-free paradigm with the key insight that the difficulty of multi-shot generation arises from a structural bias toward temporal continuity in pretrained video diffusion models, and consequently, propose a unified framework named CineWeaver to achieve reference-controllable multi-shot long-video generation without retraining. We manipulate positional encoding and attention patterns to break temporal continuity during inference to enable clear shot transitions using pretrained video diffusion models. Furthermore, we extend the proposed framework with a shot-routed reference conditioning mechanism for per-shot fine-grained controllability, and develop an anchor memory mechanism to allow long-form generation with consistent global appearance cues. To our best knowledge, CineWeaver is the first unified framework to simultaneously enable \textbf{long-form}, \textbf{reference-controllable}, and \textbf{multi-shot} video generation in a training-free fashion. Experimental results demonstrate that CineWeaver produces high-quality cinematic videos of long durations with consistent identities, stable global appearance, and clear shot transitions. The project page is available at: https://cineweaver.github.io.