arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ShotPlan:使用可学习规划令牌生成电影级视频

ShotPlan: Cinematic Video Generation with Learnable Planning Token

Su Guo, Guangce Liu, Haosen Yang, Jiepeng Wang, Cong Liu, Junqi Liu, Haibin Huang, Hongxun Yao, Chi Zhang, Xuelong Li

arXiv 2607.17675首次发表:更新:

发表机构

Institute of Artificial Intelligence (TeleAI), China Telecom; Harbin Institute of Technology(中国电信人工智能研究院(TeleAI); 哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对电影级视频生成难题,提出ShotPlan框架,基于视频扩散基础模型,引入可学习规划令牌捕捉镜头过渡线索,结合分数时间旋转位置嵌入在帧级别建模,实验证明其优于现有方法,提升了镜头管理与一致性。

AI 中文摘要

当前视频生成模型在单镜头生成方面成果显著,但在电影级视频生成中存在局限,因其需要连贯叙事和有效的多镜头构图,这需要明确的镜头规划。为应对这一挑战,我们提出ShotPlan,这是一个基于视频扩散基础模型的用于明确多镜头电影级视频生成的框架。我们的方法引入了可学习的规划令牌,它能捕捉镜头级过渡线索,并可与原始视频生成令牌无缝集成以控制过渡时间戳。与标准视频生成令牌不同,所提出的规划令牌配备了分数时间旋转位置嵌入(FRoPE),使镜头过渡能在帧级别建模。实验表明,ShotPlan显著优于现有的电影级视频生成方法,提供了更灵活的镜头管理和更强的镜头间一致性。

英文摘要

Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot planning. To address this challenge, we propose ShotPlan, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model. Our method introduces learnable planning tokens that capture shot-level transition cues and can be seamlessly integrated with the original video generation tokens to control transition timestamps. Unlike standard video generation tokens, the proposed planning tokens are equipped with Fractional Temporal Rotary Position Embedding (FRoPE), enabling shot transitions to be modeled at the frame level. Experiments demonstrate that ShotPlan significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.

CommentsProject page: https://pensioner-11.github.io/ShotPlan/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑