发表机构
Navers Lab, Einsia.AI; Peking University; Tsinghua University(Navers实验室,Einsia.AI; 北京大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频世界模型物理不一致问题,提出基于测量的基准,含40个任务覆盖多物理领域,评估8个模型,最佳得分57.76,为诊断与进展提供可解释依据。
AI 中文摘要
视频世界模型能够生成视觉上令人信服但物理上不一致的序列,这引发了对其在具身AI系统中用于预测和规划的可靠性的担忧。现有评估通常依赖基于模型的判断或参考视频,而直接物理测试主要集中于力学领域。我们引入了“世界模型物理终极考试”,这是一个基于测量的基准,用于评估视频世界模型的物理一致性。该基准包含40个受控任务,涵盖力学、光学、流体、热学与相变现象、电磁学以及表面张力。每个任务将初始图像和生成提示与预定义的物理标准配对,使得无需参考视频即可对可观察的物理关系进行可解释的测试。其评估器结合了任务可观察性筛选与任务特定的定量物理测量。对八个视频生成模型在1280个视频上的实验揭示了持续的物理不一致性和跨任务的显著差异,最佳模型总体得分为57.76分(满分100分)。对具有已知物理关系的合成视频的评估提供了在受控条件下测量模块有效性的证据。该评估器在任务内排名和成对比较中,与人类判断的一致性也高于直接视觉语言模型基线。通过结合跨物理领域的覆盖范围、基于可测量证据的分数以及明确的测量局限性,该基准为诊断物理不一致性和追踪物理一致视频世界模型的进展提供了可解释的基础。
英文摘要
Video world models can produce visually convincing yet physically inconsistent sequences, raising concerns about their reliability for prediction and planning in embodied AI systems. Existing evaluations often rely on model-based judgments or reference videos, while direct physical tests largely focus on mechanics. We introduce World Models' Last Exam in Physics, a measurement-based benchmark for evaluating physical consistency in video world models. The benchmark comprises 40 controlled tasks spanning mechanics, optics, fluids, thermal and phase-change phenomena, electromagnetism, and surface tension. Each task pairs an initial image and a generation prompt with predefined physical criteria, enabling interpretable tests of observable physical relationships without requiring reference videos. Its evaluator combines task-observability screening with task-specific quantitative physical measurements. Experiments on eight video generation models across 1,280 videos reveal persistent physical inconsistencies and substantial variation across tasks, with the best model achieving an overall score of 57.76 out of 100. Evaluation on synthetic videos with known physical relationships provides evidence for the validity of the measurement module under controlled conditions. The evaluator also achieves higher agreement with human judgments than a direct vision-language model baseline in both within-task rankings and pairwise comparisons. By combining coverage across physical domains with scores grounded in measurable evidence and explicit measurement limitations, the benchmark provides an interpretable basis for diagnosing physical inconsistencies and tracking progress toward physically consistent video world models.