发表机构
S-Lab, Nanyang Technological University; The Chinese University of Hong Kong(南洋理工大学S-Lab; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视频模型评估仅看输出层面的不足,提出Apple-PI基准,含数据集、协议和评估套件三部分,通过对11个模型测试发现其距可靠世界模拟器差距大,并揭示了一些瓶颈与差距,为指导视频模型发展提供诊断基础。
AI 中文摘要
现代视频生成模型被誉为对物理定律有内在理解的新兴世界模型。然而现有基准大多仅在输出层面评估物理合理性,未验证模型是否通过忠实的、基于定律的推理过程得出结果。我们引入了Apple-PI,首个明确以物理定律为锚点的视频模型评估基准。它由三部分组成:包含400个视频的Orchard数据集,涵盖经典力学十个规范任务,区分单定律与多定律任务;基于科学推理的三阶段基准协议,包括感知、公式化和推导,对首帧进行信息图表注释并使用帧链提示;结合基于MLLM的主观评分与基于物理定律的客观度量的混合评估套件。对11个模型的基准测试表明当前视频模型距可靠的基于定律的世界模拟器仍有很大差距,我们的分析还揭示了感知到公式化再到推导的瓶颈、多定律状态转移薄弱以及模拟到现实的持续差距。这些发现使Apple-PI成为指导未来视频模型成为具有基于法律的物理智能的世界模型的诊断基础。
英文摘要
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-PI comprises three components. 1) Orchard: a dataset of 400 videos covering ten canonical tasks in classical mechanics. It separates single-law tasks for confounder-free diagnosis from multi-law tasks for probing generalization. 2) Benchmark Protocol: a three-stage protocol based on scientific reasoning, including Perception, Formulation, and Deduction. It uses chain-of-frames prompting on infographic-annotated first frames, treating the generated video as the model's visible reasoning trace. 3) Evaluation Suite: a hybrid evaluation suite that combines MLLM-based subjective scoring with physics-law-grounded objective measures. This enables stage-resolved diagnosis of not only whether a model fails, but where it fails. Benchmarking 11 models shows that current video models remain far from reliable law-grounded world simulators, with the best video model scoring only 0.473. Our stage-, pillar-, and source-resolved analyses further expose a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. These findings position Apple-PI as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.