发表机构
Carnegie Mellon University; University of Michigan; Toyota Research Institute(卡内基梅隆大学; 密歇根大学; 丰田研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出H2RBench真实到仿真基准,标准化评估人机迁移方法,发现方法利用人类数据能力差异大,且仿真性能可预测真实性能(皮尔逊r=0.89)。
AI 中文摘要
从人类视频演示中学习机器人操作策略是规模化机器人学习的一条有前景的途径。然而,比较不同的人到机器人(H2R)迁移方法仍然具有挑战性,因为现有方法在不同的设置下进行评估,包括不同的任务套件、场景布局、物体实例和机器人监督量。为解决这一挑战,我们提出了H2RBench,一个用于评估H2R迁移方法的真实到仿真(Real2Sim)基准。H2RBench提供了一个基于真实人类视频演示和仿真机器人演示的标准化协议,并包含四个涵盖不同交互需求的操作任务。我们评估了多种具有代表性的H2R迁移方法,每种方法都采用不同的策略来弥合具身差距。利用H2RBench,我们系统地刻画了每种方法随人类演示数量增加而扩展的性能,揭示了不同方法在利用额外人类数据的能力上存在显著差异。我们进一步表明,仿真性能在广泛程度上可以预测真实世界机器人性能,在方法-任务配置上,总体皮尔逊相关系数为r=0.89,斯皮尔曼相关系数为ρ=0.85,平均最大排名违规(MMRV)为0.06。这些结果确立了H2RBench作为在实际部署前进行H2R比较评估的实用且可扩展的基准。
英文摘要
Learning robot manipulation policies from human video demonstrations constitutes a promising avenue for scalable robot learning. However, comparing different human-to-robot (H2R) transfer methods remains challenging, as existing approaches are evaluated under different settings, including differing task suites, scene layouts, object instances, and amounts of robot supervision. To address this challenge, we present H2RBench, a Real2Sim benchmark for evaluating H2R transfer methods. H2RBench provides a standardized protocol built on real human video demonstrations and simulated robot demonstrations, and includes four manipulation tasks spanning diverse interaction requirements. We evaluate multiple representative H2R transfer methods, each adopting a different strategy for bridging the embodiment gap. Using H2RBench, we systematically characterize how each method scales with the amount of human demonstrations, revealing that methods differ substantially in their ability to leverage additional human data. We further show that simulation performance is broadly predictive of real-world robot performance, with an overall Pearson correlation of r = 0.89, Spearman correlation of \r{ho} = 0.85 and Mean Maximum Rank Violation (MMRV) of 0.06 across method-task configurations. These results establish H2RBench as a practical and scalable benchmark for comparative H2R evaluation prior to real-world deployment.
Comments10th Conference on Robot Learning (CoRL 2026), Austin, TX, USA