发表机构
Institute of Artificial Intelligence (TeleAI), China Telecom; National University of Singapore; Fudan University; Tsinghua University; Shenzhen Research Institute of Northwestern Polytechnical University(中国电信人工智能研究院; 新加坡国立大学; 复旦大学; 清华大学; 西北工业大学深圳研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在评估具身世界模型物理一致性,提出无逆动力学模型的KineBench基准测试。利用级联视觉基础模型提取姿态,在物理模拟器中闭环验证,纳入运动学指标,基于ManiSkill3的任务评估EWMs,揭示任务复杂度受限的非线性缩放,为数据缩放策略提供指导。
AI 中文摘要
评估具身世界模型(EWMs)的物理一致性是一项关键的开放挑战。虽然通过模拟器展开进行闭环评估比开环评估能更忠实地评估物理合理性,但现有框架几乎完全依赖逆动力学模型(IDMs)进行动作提取。由于从2D像素空间到3D运动学空间的复杂映射,学习到的IDMs对其训练分布之外的数据可能很脆弱,导致从包含新物体和场景的生成视频中提取动作不可靠。为减少这种模糊性,我们提出了KineBench,这是一个基于显式运动学基础管道的EWMs无IDM闭环基准测试。给定一个生成的视频,KineBench使用级联视觉基础模型从单个帧中直接提取6D末端执行器姿态,然后在物理模拟器中执行以进行闭环验证。除了基于执行的任务成功之外,KineBench还纳入了两个经典的3D运动学指标——光谱弧长(SPARC)和丸山可操作性指数——从机器人中心视角表征轨迹平滑度和运动学可行性。基于ManiSkill3中的20个不同操作任务,KineBench在四个渐进套件中评估EWMs:基本执行、任务转移、视觉分布外泛化和复杂度条件缩放。对前沿模型的评估揭示了具身视频生成中任务复杂度受限的非线性缩放,为未来的数据缩放策略提供了实证指导。
英文摘要
Evaluating the physical consistency of embodied world models(EWMs) is a critical open challenge. While closed-loop evaluation via simulator rollouts offers a more faithful assessment of physical plausibility than open-loop alternatives, existing frameworks almost exclusively rely on Inverse Dynamics Models(IDMs) for action extraction. Due to the intricate mapping from 2D pixel space to 3D kinematic space, the learned IDMs can be brittle to data outside their training distribution, resulting in unreliable action extraction from the generated videos with novel objects and scenarios. This creates an unavoidable attribution ambiguity between world model inaccuracies and extractor errors. To reduce this ambiguity, we present KineBench, an IDM-free closed-loop benchmark for EWMs, built upon an explicit kinematic grounding pipeline. Given a generated video, KineBench employs cascaded visual foundation models to directly extract 6D end-effector poses from individual frames, which are then executed in a physics simulator for closed-loop validation. Beyond execution-based task success, KineBench incorporates two classical 3D kinematic metrics--Spectral Arc Length (SPARC) and the Maruyama Manipulability Index--to characterize trajectory smoothness and kinematic feasibility from a robot-centric perspective. Built on 20 diverse manipulation tasks in ManiSkill3, KineBench evaluates EWMs across four progressive suites: basic execution, task transfer, visual out-of-distribution generalization, and complexity-conditioned scaling. Evaluation across frontier models reveals task-complexity-bounded nonlinear scaling in embodied video generation, providing empirical guidance for future data-scaling strategies.
CommentsAccept to ECCV2026