LogiShot:逻辑连贯的跨镜头视频生成
LogiShot: Logically Coherent Cross-Shot Video Generation
浏览论文内容
中文总结 AI 辅助
针对跨镜头视频生成的内容脱节问题,提出LogiShot模型,通过联合编码与视觉记忆路径实现逻辑连贯,构建11万样本数据集及基准,性能优于现有方法。
中文摘要 AI 辅助
生成逻辑连贯的跨镜头视频对内容创作至关重要。当前,大多数跨镜头视频生成工作流(如短剧制作)仍依赖孤立的文本脚本或显式参考图像来指定生成内容。因此,当用户指令不明确或模糊时,生成的片段可能在视觉上看似合理,但无法与整体叙事对齐,导致内容脱节。我们认为,在视频生成中实现跨镜头逻辑连贯性需要建立镜头间的逻辑联系并保持视觉一致性。为此,我们提出LogiShot,它通过两条互补路径整合信息:1)LogiShot联合编码上下文视频和其他条件信号,生成密集多模态线索,为跨镜头生成提供视觉语义证据;2)模型在生成过程中维护上下文视频的视觉记忆,以保持镜头间的视觉一致性。此外,我们构建了包含11万个样本的数据集和专用基准,用于评估跨镜头逻辑连贯性。实验表明,LogiShot在多个镜头的逻辑连贯性方面始终优于现有基线,模型和数据将公开提供。
英文摘要
Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production, still rely on isolated textual scripts or explicit reference images to specify the generated content. Consequently, when user instructions are underspecified or ambiguous, a generated clip may appear visually plausible on its own but fail to align with the overall narrative, leading to disjointed content. We argue that achieving cross-shot logical coherence in video generation requires establishing logical connections across shots and maintaining visual consistency. To this end, we propose LogiShot, which incorporates information through two complementary paths: 1) LogiShot jointly encodes the context video and other conditioning signals, yielding dense multimodal cues that provide visual-semantic evidence for cross-shot generation; 2) the model maintains a visual memory of the context video throughout generation to preserve visual consistency across shots. Additionally, we construct a dataset with 110K samples and a dedicated benchmark for evaluating cross-shot logical coherence. Experiments demonstrate that LogiShot consistently outperforms existing baselines in terms of logical coherence across multiple shots. Model and data will be made publicly available.