AI 中文总结
该研究推出H2R-Bench基准测试,评估视频世界模型跨实体将人类操作视频转为机器人操作视频的能力,发现现有模型在实体一致性等方面存在局限,为相关模型评估提供诊断框架。
AI 中文摘要
大规模操作数据对机器人学习至关重要,但收集机器人演示的成本高昂且难以扩展。与此同时,大量以自我为中心的人类操作视频提供了丰富的行为经验,但由于人手与机器人末端执行器之间存在差异,将这些视频跨实体迁移仍然具有挑战性。视频世界模型的最新进展为从人类观测中合成以机器人为中心的操作视频提供了有前景的途径,但其跨实体迁移能力在很大程度上仍未被探索。因此,我们推出H2R-Bench,这是一个用于评估跨实体人到机器人操作视频生成的基准测试,其中模型将以自我为中心的人类演示转换为特定实体下的机器人操作视频。每个基准测试实例包含人类演示视频、目标实体约束以及涵盖任务目标、动作事件、功能接触和对象响应的源标注。H2R-Bench通过五个维度评估生成的视频,包括目标状态完成度、动作事件完成度、功能接触迁移、实体正确性和通用视频质量。我们在六个操作类别和两个机器人实体上对十一种最先进的视频生成模型进行了基准测试。我们的评估显示,当前的视频世界模型在人到机器人操作迁移方面仍然存在局限:即使是领先的模型也常常在实体一致性、功能交互和任务执行方面失败。H2R-Bench提供了一个系统的诊断框架,用于评估视频世界模型是否能够弥合人到机器人的实体差距,并将人类操作观测转换为以机器人为中心的训练资源。
英文摘要
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.