发表机构
Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视频生成模型忽视物理属性导致视频不真实的问题,提出PhysicsLENS数据集与基准,通过匹配场景对评估七个物理域,发现模型常忽略属性且合理性下降不显著。
AI 中文摘要
可靠的视频世界模型可以为机器人学习、规划和评估提供可扩展的预测环境。然而,生成的机器人视频可能违反物理原理,并通过物理上不切实际的行为完成任务,这限制了它们在机器人学习和规划中的可靠性。当前的视频生成基准排除了本质上被视觉隐藏的物理属性(例如重量、粘性、摩擦)。因此,视频模型在物理保真度上被评估,而非物理的潜在准确性。我们引入了PhysicsLENS,一个用于评估基于机器人的物理属性合理性的数据集和基准。PhysicsLENS使用匹配的场景对,这些场景对保持相同的条件帧和任务,同时在场景描述中改变底层物理属性。场景从公共机器人视频源中精选,并在七个物理领域进行标注:碰撞、重力、动量、摩擦、变形、流体和因果性。我们评估了四种视频生成模型,产生了超过400个人工标注标签。结果表明,看似合理的视频常常忽略所述属性(47例中有34例),并且陈述属性仅略微降低合理性,且不显著。
英文摘要
Reliable video world models could provide scalable predictive environments for robot learning, planning, and evaluation. However, generated robot videos can violate physical principles and complete tasks through physically implausible behavior, limiting their reliability for robot learning and planning. Current video-generation benchmarks exclude physics that are inherently hidden by visuals (e.g., weight, viscosity, friction). Due to this, video models are evaluated on the fidelity of physics, not the underlying accuracy of physics. We introduce PhysicsLENS, a dataset and benchmark for evaluating plausibility of physical properties grounded in robotics. PhysicsLENS uses matched scenario pairs that hold the same conditioning frame and task, while varying underlying physics in the scene description. Scenarios are curated from public robot video sources and annotated across seven physical domains: collision, gravity, momentum, friction, deformation, fluid, and causality. We evaluate across four video generation models, producing over 400 human-annotated labels. Results show that plausible-looking videos often ignore the stated property (34 of 47), and that stating the property lowers plausibility only slightly and not significantly.