VideoGen-of-Thought:以最少人工干预逐步生成多镜头视频
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
- Hong Kong University of Science and Technology(香港科技大学)
- Peking University(北京大学)
- University of Hong Kong(香港大学)
- NUS(南洋理工大学)
- University of Central Florida(佛罗里达大学)
- Everlyn AI
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对现有视频生成模型多镜头叙事断裂、视觉不一致等问题,提出无需训练的VGoT框架,通过动态叙事建模、身份感知跨镜头传播和相邻潜空间转场,以极少人工实现高质量多镜头视频生成。
中文摘要 AI 辅助
当前的视频生成模型擅长生成短片段,但由于视觉动态不连贯、故事情节断裂,无法生成具有凝聚力的多镜头叙事内容。现有解决方案要么依赖大量的人工脚本编写/编辑,要么优先考虑单镜头保真度而非跨场景连续性,限制了其在类电影内容制作中的实用性。我们提出了VideoGen-of-Thought(VGoT,视频生成思维链),这是一个逐步式框架,仅需单句输入即可自动完成多镜头视频合成,系统性地解决了三大核心挑战:(1)叙事碎片化:现有方法缺乏结构化的叙事设计。我们提出动态故事情节建模,将用户提示转化为简洁的镜头草稿,再扩展为五个维度的详细规范(角色动态、背景连续性、关系演变、镜头运动和HDR打光),并通过自验证确保逻辑推进。(2)视觉不一致性:以往方法难以在不同镜头间保持外观一致性。我们的身份感知跨镜头传播机制构建了身份保留肖像(identity-preserving portrait,IPP)令牌,在维持角色身份的同时,允许剧情所需的可控特征变化(表情、年龄增长)。(3)转场伪影:镜头切换突兀会破坏沉浸感。我们的相邻潜空间转场机制采用边界感知重置策略,在转场点处理相邻镜头的特征,在保留叙事连续性的同时实现流畅的视觉过渡。在无需训练的流水线中整合上述模块后,VGoT的镜头内人脸一致性比强基线模型高出20.4%,风格一致性高出17.4%,同时所需人工调整量减少10倍。VGoT为自动化多镜头视频生成填补了原始视觉合成与导演级叙事之间的空白。
英文摘要
Current video generation models excel at short clips but fail to produce cohesive multi-shot narratives due to disjointed visual dynamics and fractured storylines. Existing solutions either rely on extensive manual scripting/editing or prioritize single-shot fidelity over cross-scene continuity, limiting their practicality for movie-like content. We introduce VideoGen-of-Thought (VGoT), a step-by-step framework that automates multi-shot video synthesis from a single sentence by systematically addressing three core challenges: (1) Narrative fragmentation: Existing methods lack structured storytelling. We propose dynamic storyline modeling, which turns the user prompt into concise shot drafts and then expands them into detailed specifications across five domains (character dynamics, background continuity, relationship evolution, camera movements, and HDR lighting) with self-validation to ensure logical progress. (2) Visual inconsistency: previous approaches struggle to maintain consistent appearance across shots. Our identity-aware cross-shot propagation builds identity-preserving portrait (IPP) tokens that keep character identity while allowing controlled trait changes (expressions, aging) required by the story. (3) Transition artifacts: Abrupt shot changes disrupt immersion. Our adjacent latent transition mechanisms implement boundary-aware reset strategies that process adjacent shots' features at transition points, enabling seamless visual flow while preserving narrative continuity. Combined in a training-free pipeline, VGoT surpasses strong baselines by 20.4\% in within-shot face consistency and 17.4\% in style consistency, while requiring 10x fewer manual adjustments. VGoT bridges the gap between raw visual synthesis and director-level storytelling for automated multi-shot video generation.