互补检索增强提示用于一致的长视频生成
Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation
浏览论文内容
中文总结 AI 辅助
提出互补检索增强提示框架,通过解析脚本构建视觉元素注册表并检索互补历史参考,实现无需训练的长视频一致生成,优于现有检索基线。
中文摘要 AI 辅助
虽然最近的视频基础模型在生成高质量短视频方面表现出色,但长视频生成仍然是一个关键挑战,其中一个主要瓶颈在于如何对独立生成的镜头进行条件约束,以在整个故事中保持角色、场景和物体的一致性。现有的免训练方法通常使用检索到的历史视觉内容来对目标镜头进行条件约束。然而,这些参考往往存在严重的信息不匹配,要么引入无关的上下文冗余,要么未能提供目标镜头所需元素的完整组合。为解决这一问题,我们提出了互补检索增强提示(Complementary Retrieval-Augmented Prompting),这是一种智能体框架,通过策略性地聚合一组紧凑且相互支持的历史参考,实现对长视频生成的完整且有针对性的条件约束,而无需重新训练或修改底层生成器。具体来说,我们的框架通过将叙事脚本解析为基于文本的视觉元素注册表来显式建模每个目标镜头所需的视觉元素,该注册表跟踪角色、物体、场景及其镜头级状态。一个由视觉语言模型(VLM)标注的关键帧库进一步将这些元素映射到过去的视觉观察。在所需元素的引导下,我们的智能体检索互补参考,以最大化目标元素覆盖率,同时最小化历史噪声。最后,检索到的参考、结构化元素状态和基础指令被组装成一个统一的提示,输入到冻结的视频生成器中。这种元素感知的过程提供了全面的条件约束,同时保持完全可解释。在多镜头故事生成的定量和定性评估中,我们的方法在跨镜头一致性和文本可控性方面始终优于基于最近帧、基于记忆和基于实体级别的检索基线。
英文摘要
While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free approaches typically condition target shots using retrieved historical visuals. However, these references often suffer from severe informational mismatch, either introducing irrelevant contextual redundancy or failing to provide the full combination of required elements for the target shot. To resolve this, we present Complementary Retrieval-Augmented Prompting, an agentic framework that strategically aggregates a compact set of mutually supportive historical references to achieve complete and targeted conditioning for long-form video generation without retraining or modifying the underlying generator. Specifically, our framework explicitly models the visual elements required by each target shot by parsing the narrative script into a text-grounded visual element registry that tracks characters, objects, scenes, and their shot-level states. A VLM-annotated keyframe library further maps these elements to past visual observations. Guided by the required elements, our agent retrieves complementary references that maximize target-element coverage while minimizing historical noise. Finally, the retrieved references, structured element states, and grounding instructions are assembled into a unified prompt for the frozen video generator. This element-aware process provides comprehensive conditioning while remaining fully interpretable. Quantitative and qualitative evaluations on multi-shot story generation demonstrate that our method consistently outperforms recent-frame, memory-based, and entity-level retrieval baselines in cross-shot consistency and text-controllability.