arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2503.15138cs.CV

VideoGen-of-Thought:以最少人工干预逐步生成多镜头视频

VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention

Mingzhe Zheng, Yongqi Xu, Haojian Huang, Xuran Ma, Yexin Liu, Wenjie Shu, Yatian Pang, Feilong Tang, Qifeng Chen, Harry Yang, Ser-Nam Lim

更新

AI总结:

针对现有视频生成模型多镜头叙事碎片化、视觉不一致、转场有伪影的问题,提出VGoT逐步框架,通过动态叙事建模、身份感知跨镜头传播、相邻潜变量转场三大机制,实现单句提示自动生成高质量多镜头视频,性能显著优于基线。

AI中文摘要:

当前视频生成模型擅长生成短片段,但因视觉动态不连贯和故事情节碎片化,无法生成具有凝聚力的多镜头叙事内容。现有解决方案要么依赖大量人工脚本编写/编辑,要么优先考虑单镜头保真度而非跨场景连续性,限制了其在类电影内容创作中的实用性。我们提出了VideoGen-of-Thought(VGoT),这是一个逐步式框架,可通过系统解决三大核心挑战,仅凭单个句子自动完成多镜头视频合成:(1)叙事碎片化:现有方法缺乏结构化叙事能力。我们提出动态故事情节建模,先将用户提示词转换为简洁的镜头描述,再在五个维度(角色动态、背景连续性、关系演变、镜头运动、HDR打光)将其细化为详细的电影级规格,并通过自验证确保叙事逻辑推进。(2)视觉不一致性:现有方法难以维持跨镜头的视觉一致性。我们的身份感知跨镜头传播机制会生成保身份肖像(identity-preserving portrait, IPP)token,在维持角色保真度的同时,允许其呈现故事情节要求的特征变化(表情、年龄增长)。(3)转场伪影:镜头切换突兀会破坏沉浸感。我们的相邻潜变量转场机制采用边界感知重置策略,在转场点处理相邻镜头的特征,在保留叙事连续性的同时实现流畅的视觉过渡。VGoT生成的多镜头视频在镜头内人脸一致性上优于现有最优基线20.4%,风格一致性提升17.4%,同时跨镜头一致性提升超100%,所需人工调整量比其他方法少10倍。

英文摘要:

Current video generation models excel at short clips but fail to produce cohesive multi-shot narratives due to disjointed visual dynamics and fractured storylines. Existing solutions either rely on extensive manual scripting/editing or prioritize single-shot fidelity over cross-scene continuity, limiting their practicality for movie-like content. We introduce VideoGen-of-Thought (VGoT), a step-by-step framework that automates multi-shot video synthesis from a single sentence by systematically addressing three core challenges: (1) Narrative Fragmentation: Existing methods lack structured storytelling. We propose dynamic storyline modeling, which first converts the user prompt into concise shot descriptions, then elaborates them into detailed, cinematic specifications across five domains (character dynamics, background continuity, relationship evolution, camera movements, HDR lighting), ensuring logical narrative progression with self-validation. (2) Visual Inconsistency: Existing approaches struggle with maintaining visual consistency across shots. Our identity-aware cross-shot propagation generates identity-preserving portrait (IPP) tokens that maintain character fidelity while allowing trait variations (expressions, aging) dictated by the storyline. (3) Transition Artifacts: Abrupt shot changes disrupt immersion. Our adjacent latent transition mechanisms implement boundary-aware reset strategies that process adjacent shots' features at transition points, enabling seamless visual flow while preserving narrative continuity. VGoT generates multi-shot videos that outperform state-of-the-art baselines by 20.4% in within-shot face consistency and 17.4% in style consistency, while achieving over 100% better cross-shot consistency and 10x fewer manual adjustments than alternatives.

补充信息

↑