发表机构
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; EvoPhys AI; Joy Future Academy, JD; The University of Sydney; The Hong Kong University of Science and Technology; Beijing Institute of Technology(北京大学计算机学院多媒体信息处理国家重点实验室; EvoPhys人工智能公司; 京东探索研究院; 悉尼大学; 香港科技大学; 北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出WorldSimProbe框架,通过五套受控测试评估动作条件世界模型的模拟器保真度,发现其存在动作实现退化等问题,为具身操控模型提供标准化诊断方法。
AI 中文摘要
动作条件世界模型(ACWMs)有望为具身AI提供可扩展的预测模拟器,用于规划、策略评估和数据生成。要实现这一承诺,需要精确的动作条件转换,而非仅合理的输出。然而,其适用性难以确定,因为主流评估强调视觉质量、任务结果或粗略的滚动级响应,未直接测试模拟器保真度。为解决这一差距,我们通过物理模拟器应具备的可观测能力来评估ACWMs,据此将「可观测模拟器契约」形式化,这是任何动作条件物理模拟器都应满足的最小契约:输入动作必须诱导对应的智能体运动,环境响应必须基于该已实现的运动。为实施该契约,我们引入WorldSimProbe,包含五个受控套件,涵盖局部控制敏感性、全局轨迹变化、源多样化动作、交互基础和动力学。特定套件的评估器评估模拟器相对校准、密集动作-运动对应关系、虚假交互基础和原语级动力学。我们在RoboTwin、ManiSkill和LIBERO上对六个开源ACWMs的超过18000个实例进行评估。WorldSimProbe揭示了控制变化下的系统性动作实现退化、交互基础和动力学中的结构化故障,以及与人类判断和下游结果一致的基准信号。总之,这个基于能力的框架提供了一种透明且标准化的范式,用于诊断ACWM的模拟器保真度,超越了粗略的任务导向评估。
英文摘要
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this promise requires precise action-conditioned transitions rather than merely plausible outputs. Yet their applicability remains difficult to establish because prevailing evaluations emphasize visual quality, task outcomes, or coarse rollout-level responsiveness without directly testing simulator fidelity. To address this gap, we evaluate ACWMs through the observable capabilities expected of physical simulators. Accordingly, we formalize Observable Simulator Contract, a minimal contract that any action-conditioned physical simulator should satisfy: supplied actions must induce corresponding agent motion, and environment responses must be grounded in that realized motion. To operationalize this contract, we introduce WorldSimProbe, comprising five controlled suites spanning local control sensitivity, global trajectory variation, source-diverse actions, interaction grounding, and dynamics. Suite-specific evaluators assess simulator-relative calibration, dense action-to-motion correspondence, false-interaction grounding, and primitive-level dynamics. We evaluate six open-source ACWMs on more than 18,000 instances across RoboTwin, ManiSkill, and LIBERO. World-SimProbe reveals systematic action-realization degradation across control variation, structured failures in interaction grounding and dynamics, and benchmark signals consistent with human judgments and downstream outcomes. Together, this capability-based framework provides a transparent, and standardized paradigm for diagnosing ACWM simulator fidelity beyond coarse, task-directed evaluation.
Comments20 pages, 18 figures, and 10 tables, including supplementary material. Code and data: https://evophys.com/WorldSimProbe/