arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StoryEngine:一种基于状态锚定的智能体视频叙事框架

StoryEngine: A State-Grounded Agentic Framework for Video Storytelling

Yingrui Wang, Zeqing Wang, Yeying Jin

arXiv 2609.33627首次发表:更新:

发表机构

Tencent; National University of Singapore(腾讯; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

StoryEngine提出基于状态锚定的智能体框架,通过分离语义计划与视觉观察、传播事件状态并引入修复循环,实现连贯一致的长篇视频叙事,实验全面优于现有方法。

AI 中文摘要

尽管智能体驱动的多镜头视频生成近期取得了进展,但生成连贯且一致的长篇故事仍然具有挑战性。现有的智能体流程通常依赖于文本镜头计划或先前生成的像素,但缺乏一种显式机制来传播故事事件的后果并在各镜头间维护视频世界状态。因此,缺失的视觉细节可能被不准确地重建,而视觉漂移可能在后续镜头中传播,从而削弱叙事连贯性和视觉一致性。为应对这些挑战,我们提出了StoryEngine,一种基于状态锚定的智能体视频叙事框架。StoryEngine在权威语义计划与不可靠视觉观察之间建立了分离。具体而言,StoryEngine维护实体位置和故事相关状态的结构化表示,并传播事件引发的变化以定义每个镜头的预期起始和结束状态。为在视觉上实现这些状态,StoryEngine为重复出现的实体和环境构建规范参考,并将状态和视觉约束编译为可执行的渲染计划。同时,为正确实现这些状态,一个有界评估引导的修复循环进一步纠正局部状态不一致。这些机制共同保持了因果故事进展,并防止局部视觉错误在镜头间传播。为全面评估长篇叙事,我们构建了一个涵盖多样场景和视觉风格的基准,并采用评估叙事质量、叙事连贯性和视觉一致性的指标。实验结果表明,StoryEngine在所有评估维度上均持续优于最先进方法,验证了其在连贯且一致的视频叙事中的有效性。

英文摘要

Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, undermining both narrative coherence and visual consistency. To address these challenges, we propose StoryEngine, a state-grounded agentic framework for video storytelling. StoryEngine establishes a separation between authoritative semantic plans and unreliable visual observations. Specifically, StoryEngine maintains a structured representation of entity placement and story-relevant states, and propagates event-induced changes to define the intended start and end states of each shot. To visually realize these states, StoryEngine constructs canonical references for recurring entities and environments, and compiles state and visual constraints into executable render plans. Meanwhile, to realize these states correctly, a bounded evaluation-guided repair loop further corrects local state inconsistencies. Together, these mechanisms preserve causal story progression and prevent local visual errors from propagating across shots. To comprehensively evaluate long-form storytelling, we construct a benchmark across diverse scenarios and visual styles, with metrics assessing storytelling quality, narrative coherence, and visual consistency. Experimental results demonstrate that StoryEngine consistently outperforms state-of-the-art methods across all evaluation dimensions, validating its effectiveness for coherent and consistent video storytelling.

Comments23 pages, 4 figures, submit to ICLR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑