发表机构
Peking University; Tsinghua University; Beihang University; Shanghai Jiao Tong University; University of Science and Technology of China; Shanghai AI Laboratory(北京大学; 清华大学; 北京航空航天大学; 上海交通大学; 中国科学技术大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TriWorldBench通过同步头、左腕、右腕视频评估具身世界模型的三视角一致性,含500片段、19指标,扩展评估至多视角预测一致性。
AI 中文摘要
具身世界模型预测机器人动作的结果,以支持学习和规划。对于配备头部和腕部摄像头的机器人,这需要互补的视角:头部视角捕捉整体任务,而腕部视角揭示局部夹爪-物体交互。然而,独立评估这些视角无法确定它们是否描述相同的动作和物体状态。我们引入了TRIWORLDBENCH,一个通过同步的头部、左腕和右腕视频来评估具身世界模型的基准。它包含50个双臂操作任务中的500个片段,并使用19个指标来评估三视角一致性、任务对齐、物理和3D连贯性、运动质量、时间一致性和视觉质量。通过将跨视角检查与针对每个摄像头定制的测量相结合,该基准评估了看似合理的单个视频是否也构成对预期任务的一致预测。我们用TWB-Score总结整体性能,并保留每个视角的结果以识别预测失败之处。这将世界模型评估扩展到单视角视觉质量之外。代码、数据和指标定义可在https URL获取。
英文摘要
Embodied world models predict the outcomes of robot actions to support learning and planning. For robots equipped with head and wrist cameras, this requires complementary views: the head view captures the overall task, while wrist views reveal local gripper-object interactions. However, evaluating these views independently cannot determine whether they describe the same action and object state. We introduce TRIWORLDBENCH, a benchmark for evaluating embodied world models through synchronized head, left-wrist, and right-wrist videos. It contains 500 episodes across 50 bimanual manipulation tasks and uses 19 metrics to assess tri-view consistency, task alignment, physical and 3D coherence, motion quality, temporal consistency, and visual quality. By combining cross-view checks with measurements tailored to each camera, the benchmark evaluates whether plausible individual videos also form a consistent prediction of the intended task. We summarize overall performance with TWB-Score and retain per-view results to identify where predictions fail. This extends world-model evaluation beyond single-view visual quality. Code, data, and metric definitions are available at https://github.com/TriWorldBench/TriWorldBench.