发表机构
Peking University; Advanced Institute of Information Technology, Peking University(北京大学; 北京大学先进信息技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人操作基准缺乏物理条件多样性的问题,提出 RoboTwin-Phys 基准,连续变化 13 个物理属性并发布 5000+ 演示,揭示现有模型在物理变化下的鲁棒性差距。
AI 中文摘要
当前机器人操作基准大多缺乏物理条件多样性。虽然大规模仿真基准越来越多地融入物体外观、场景布局和视觉观察的变化,但它们通常保持底层物理参数固定不变。因此,诸如质量、摩擦和关节动力学等真实世界变化的重要来源在很大程度上仍未得到测试。我们引入了 RoboTwin-Phys,一个物理多样性基准,将物理条件多样性视为机器人操作评估的一个明确维度。该基准在物理合理范围内连续变化 13 个物理属性,提供了一个统一的设置,用于评估策略在不同物理操作条件下的表现。我们进一步发布了超过 5,000 个带有真实物理参数的专业演示,支持物理属性估计、条件感知建模和物理条件策略训练。对代表性 WAMs 和 VLAs 的评估揭示了一个显著的鲁棒性差距:在现有视觉和布局随机化下有效的模型,在物理条件变化下可能显著退化。RoboTwin-Phys 提供了系统测量和提高机器人操作中物理条件多样性鲁棒性所需的基准、数据和评估协议。
英文摘要
Physical-condition diversity is largely missing from current benchmarks for robot manipulation. While large-scale simulation benchmarks increasingly incorporate variations in object appearance, scene layout, and visual observations, they typically keep the underlying physical parameters fixed. As a result, important sources of real-world variability, such as changes in mass, friction, and joint dynamics, remain largely untested. We introduce RoboTwin-Phys, a physics-diverse benchmark that treats physical-condition diversity as an explicit dimension of robot manipulation evaluation. The benchmark continuously varies 13 physical attributes within physically plausible ranges, providing a unified setting for evaluating policies across diverse physical operating conditions. We further release more than 5,000 expert demonstrations with ground-truth physical parameters, enabling physical-attribute estimation, condition-aware modeling, and physics-conditioned policy training. Evaluations of representative WAMs and VLAs reveal a substantial robustness gap: models that remain effective under existing visual and layout randomization can degrade markedly under changes in physical conditions. RoboTwin-Phys provides the benchmark, data, and evaluation protocol needed to systematically measure and improve robustness to physical-condition diversity in robot manipulation.
Commentstechnical report for a benchmark