arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02089cs.ROcs.AI

HumanoidToolBench:从选择到移动执行的人形机器人工具使用基准

HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution

Kyochul Jang, Seohyeon Park, Ohchul Kwon, Sangjun Park, Junhyeok Choi, Seungyeop Yi, Chaeyun Kim, Sangkyu Lee, Idan Szpektor, Avi Caciularu, Jongmin Park, Youngjae Yu

首次发表
浏览论文内容

中文总结 AI 辅助

针对人形机器人工具使用缺乏联合评估的问题,提出HumanoidToolBench基准和ToolBook数据集,评估发现工具选择与任务完成存在显著差距。

中文摘要 AI 辅助

随着机器人硬件和学习方法的进步,人形机器人需要工具来执行超出其固有物理极限的任务。成功的工具使用需要选择合适的工具,并协调操作,以及在需要时进行移动以完成任务。现有基准未能在人形机器人上联合评估这些能力。我们引入了HumanoidToolBench,一个包含18个任务的基准,涵盖三个场景、三个执行级别和两种工具集模式,以及ToolBook,一个在仿真和真实Unitree G1上收集的3.1k演示数据集。对七种策略在仿真中以及三种策略在真实机器人上的评估揭示了选择合适的工具与完成任务之间的显著差距。针对GR00T N1.7的聚焦探针显示,在未见过的工具上选择准确率降低,并且在无关指令下继续执行任务。代码和数据可在以下网址获取:此https URL。

英文摘要

As robotic hardware and learning methods advance, humanoids need tools to perform tasks beyond their inherent physical limits. Successful tool use requires selecting a suitable tool and coordinating manipulation and, when needed, locomotion to complete the task. Existing benchmarks do not jointly evaluate these capabilities on a humanoid. We introduce HumanoidToolBench, an 18-task benchmark spanning three scenarios, three execution levels, and two tool-set modes, together with ToolBook, a dataset of 3.1k demonstrations collected in simulation and on a real Unitree G1. Evaluation of seven policies in simulation and three on the real robot reveals substantial gaps between selecting a suitable tool and completing the task. Focused GR00T N1.7 probes show reduced selection accuracy on unseen tools and continued task execution under unrelated instructions. Code and data are available at https://snu-pi.github.io/HumanoidToolBench/.

发表机构

  • Seoul National University(首尔大学)
  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
  • Google Research(谷歌研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑