arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04200cs.CV

Principia:面向视频模型的关系物理测试基准

Principia: Relational Physics Tests for Video Models

  • Indian Institute of Science(印度科学学院)
  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad

AI总结:

Principia是评估视频模型物理推理的校准无关基准,涵盖八种物理现象,实验显示现有视频生成器和视觉语言模型在该基准上表现远差于通用视频基准VBench。

AI中文摘要:

评估视频模型的物理推理能力十分困难,因为绝对运动测量依赖于帧率、物体尺度和相机校准,而这些在生成视频中往往是模糊或不可用的。我们提出了一种不同的方法:当同一场景中的两个物体遵循相同物理定律时,它们的运动必须满足可预测的关系,且这些关系与校准无关。我们引入了Principia,这是一个通过配对物体间的关系一致性来评估牛顿物理的基准。Principia涵盖八种物理现象——重力、恢复系数、摩擦力、转动惯量、抛体运动、动量、单摆和质量-弹簧振荡,涉及平动、转动、碰撞和振荡动力学,使用在受控协议下录制的真实场景。我们还引入了一种与校准无关的一致性分数,可直接在图像空间中量化物理违反情况。在来自六个最先进视频生成器的数千次生成结果中,所有模型在VBench上的得分约为0.8,但在Principia上的得分均未超过0.42。我们还评估了视觉语言模型检测关系物理违反情况的能力,最佳模型仅达到67%的准确率,多数模型表现接近随机水平。

英文摘要:

Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena - gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation - across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench. Vision-language models are evaluated on their ability to detect relational physics violations, with the best model achieving only 67% accuracy and most performing near chance level.

补充信息

↑