事件边界处的预见:视频世界模型中的物理预测评估
Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models
浏览论文内容
中文总结 AI 辅助
本文提出事件锚定评估方法,用真实自由落体数据检验视频世界模型预测物理后果的能力,发现模型在后果生成、时间锚定和运动实现上存在不同挑战。
中文摘要 AI 辅助
视频世界模型在很大程度上被视为物理世界的预测模型,因此被期望能够预测所观察事件的后果。然而,现有评估主要集中于参考相似性、物理定律一致性或判断合理性,仅间接估计预测能力。我们直接解决这一问题:当一次释放或撞击刚刚发生但其后果被隐藏时,世界模型能否预测接下来应该发生什么?我们引入了一种基于事件锚定的评估方法,该方法使用62段受控的真实世界自由落体录音和124个片段,涵盖三种物体类型,并带有细粒度的释放和撞击标注以及真实轨迹。该协议将后果生成、时间定位和物理实现分开评估。在六种当代视频生成和世界模型中,Runway和Veo以超过93%的比例生成释放及后续撞击事件,但往往明显延迟启动这些事件,而Cosmos-Predict-2.5和MAGI-1则经常保留事件前状态,产生很少或没有可测量的后果。在可测量的下落中,看似合理的时间并不一定意味着物理上一致的运动。我们进一步进行了一项15名参与者、20种条件的人类研究,参与者从单个事件锚定帧描述预期后果并绘制其轨迹。人类预测在总体上偏向记录的未来,同时揭示了合理延续之间的真实模糊性。总体而言,物理预见表现为一系列不同的挑战:启动后果、将其锚定在时间中以及实现其运动。
英文摘要
Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physical-law consistency, or judgment plausibility, estimating anticipation only indirectly. We address this directly: when a release or impact has just occurred but its consequence is withheld, can a world model anticipate what should happen next? We introduce an event-anchored evaluation based on 62 controlled real-world free-fall recordings and 124 clips spanning three object types, with fine-grained release and impact annotations and ground-truth trajectories. The protocol separates consequence production, temporal placement, and physical realization. Across six contemporary video generation and world models, Runway and Veo produce release and subsequent impact events at rates above 93% but often initiate them substantially late, whereas Cosmos-Predict-2.5 and MAGI-1 frequently preserve the pre-event state and produce little or no measurable consequence. Among measurable falls, plausible timing does not necessarily imply physically consistent motion. We further conduct a 15-participant, 20-condition human study in which participants describe the expected consequence from a single event-anchored frame and draw its trajectory. Human predictions favor the recorded future in aggregate while revealing genuine ambiguity among plausible continuations. Overall, physical foresight emerges as a sequence of distinct challenges: initiating a consequence, anchoring it in time, and realizing its motion.
发表机构
- National Autonomous University of Mexico(墨西哥国立自治大学)
- University of Texas at El Paso(德克萨斯大学埃尔帕索分校)
机构由 AI 辅助整理,请以论文原文为准。