arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboPhys-3D:通过三维重建进行全面的具身世界模型评估

RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

Tianyi Wang, Jiazhou Chen, Yiming Xu, Xiangyu Li, Tianyi Zeng, Chih-Hsien Chou, Ning Lu, Liang Peng, Junfeng Jiao, Christian Claudel

arXiv 2608.28718首次发表:更新:

发表机构

The University of Texas at Austin; Purdue University; Futurewei Technologies, Inc(德克萨斯大学奥斯汀分校; 普渡大学; 星纪时代科技公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出基于RoboTwin 2.0的三维具身世界模型基准RoboPhys-3D,含50个操作任务等数据,提出两个评估分数,发现Cosmos 3的RoboPhyscore最高,该分数与人类评估一致性强。

AI 中文摘要

视频世界模型越来越多地为具身AI充当数据引擎、动作规划器和模拟器,但传统的具身世界模型(EWM)基准缺乏统一的三维基准协议,无法确定生成的滚动序列是否保留了底层三维场景状态,或是否能转化为可执行动作。我们推出RoboPhys-3D,这是一个基于RoboTwin 2.0的三维基准EWM基准,涵盖四个场景下的50个操作任务,包含5000个回合和25000个多视图真实视频。RoboPhys-3D的一个核心特点是,生成的视频和真实视频通过同一三维重建管道处理,从而能区分重建诱导误差和生成诱导误差。该基准将50个互补指标组织为四个层面的18个子维度:像素级保真度、三维几何一致性、状态级理解和任务级完整性。我们还推出了Average Full Score,即对所有50个指标取平均的分层综合评估分数,以及RoboPhyscore,即对与任务成功相关性最强的指标取平均的紧凑任务对齐分数。在四个代表性视频世界模型中,Cosmos 3取得了最高的RoboPhyscore(0.6330,为真实值的92.7%),而基于状态和执行的基准指标揭示了感知和视觉语言模型判断未能捕捉到的大量失败。RoboPhyscore还与人类评估表现出强一致性(皮尔逊相关系数r=0.9761,斯皮尔曼秩相关系数ρ=0.8962),证明了基于基准、感知执行的评估对EWM能力的重要性。

英文摘要

Video world models increasingly serve as data engines, action planners, and simulators for embodied AI, but conventional embodied world model (EWM) benchmarks lack a unified 3D-grounded protocol for establishing whether generated rollouts preserve the underlying 3D scene state or translate into executable actions. We introduce RoboPhys-3D, a 3D-grounded EWM benchmark built on RoboTwin 2.0, covering 50 manipulation tasks across four regimes, with 5,000 episodes and 25,000 multi-view ground-truth videos. A defining feature of RoboPhys-3D is that generated and ground-truth videos are processed through the same 3D reconstruction pipeline, enabling reconstruction-induced error to be distinguished from generation-induced error. The RoboPhys-3D benchmark organizes 50 complementary metrics into 18 sub-dimensions across four levels: pixel-level fidelity, 3D geometry consistency, state-level understanding, and task-level completeness. We further introduce Average Full Score, a hierarchical score averaging all 50 metrics for comprehensive evaluation, and RoboPhyscore, a compact task-aligned score averaging the metrics most strongly correlated with task success. Among the four representative video world models, Cosmos 3 achieves the highest RoboPhyscore (0.6330, 92.7% of ground truth), while state- and execution-grounded metrics reveal substantial failures that perceptual and vision-language model-based judgments fail to capture. RoboPhyscore further exhibits strong agreement with human evaluation (Pearson r = 0.9761 and Spearman \r{ho} = 0.8962), demonstrating the importance of grounded, execution-aware evaluation for EWM capability.

Comments66 pages, 12 figures, 55 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑