arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03673cs.CV

视频生成器是否会跨片段追踪世界?视频延续中世界状态推理的基准与方法

Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation

  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Yingmao Miao, Pengfei Zhang, Chaoran Xu, Meng Yu, Jing Tang, Xiangxiang Chu, Chao Shen, Chenhao Lin

AI总结:

该研究针对视频生成器无法可靠追踪视频世界状态的问题,构建了基准Statebench并提出方法Stateagent,将可控视频延续的全案例状态分数从45.2提升至69.3,还有益于一分钟规模的故事生成。

AI中文摘要:

视频生成器通过拼接较短片段来构建长视频,可逐段生成或自回归扩展片段,每个新片段通常依赖历史观测的记忆,如近期帧、选定关键帧、记忆库或缓存特征。这些记忆保留了过去的可见证据,但当前生成器无法可靠地将此类证据转化为世界状态接口:即视频世界在之前动作后处于何种状态,以及在新提示下应如何变化。过去的帧仍是有效历史,但可能无法描述下一片段所需的状态;部分状态必须从被遮挡或隐含的变化中推断,而非从直接观测的帧中复制。这为视频延续提出了一个简单却被忽视的问题:给定前一段视频、其提示及新提示,模型能否生成反映历史视频与新提示共同决定的状态的延续内容?为回答该问题,我们引入Statebench,这一基准通过测试三类状态下的延续内容来填补上述缺口:过去可见状态、被遮挡过程状态及复杂转换状态。我们还提出Stateagent,它明确维护实体-状态表示,在新提示下更新该表示,将预测的动作后状态作为未来结束帧进行接地,再渲染下一段视频。实验表明,我们的方法提升了可控视频延续性能,将全案例状态分数(SCS-All)从45.2提高至69.3,同时也有益于一分钟规模的故事生成。代码可在this https URL获取。

英文摘要:

Video generators build long videos by composing shorter parts, either by generating segments one after another or by autoregressively extending chunks. Each new part usually depends on memories of historical observations, such as recent frames, selected key frames, memory banks, or cached features. These memories preserve visible evidence from the past, but current generators do not reliably turn such evidence into a world-state interface: what holds in the video world after previous actions and how it should change under the next prompt. A past frame remains valid history, but it may not describe the state needed by the next segment; some states must instead be inferred from occluded or implicit changes rather than copied from a directly observed frame. This creates a simple but overlooked question for video continuation: given a previous video, its prompt, and a new prompt, can a model generate a continuation that reflects the state determined by both the historical video and the new prompt? To answer this question, we introduce Statebench, a benchmark that targets this gap by testing continuations over three state categories: past-visible states, occluded-process states, and complex-transition states. We further propose Stateagent, which explicitly maintains an entity-state representation, updates it under the new prompt, grounds the predicted post-action state as a future end frame, and renders the next video. Experiments show that our method improves controlled video continuation by raising the all-case state score (SCS-All) from 45.2 to 69.3, and also benefits story generation at the one-minute scale. Code is avaliable at https://github.com/AMAP-ML/StateAgent.

↑