arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HRIBench:以交互为中心的人机协作基准测试

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration

Chang Liu, Jiawei Zhang, Tao Zhang, Ye Wang, Hongyu Zhou, Qin Jin

arXiv 2607.13056首次发表:更新:

发表机构

Renmin University of China; Beijing Normal University(中国人民大学; 北京师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对当前VLA基准未充分考量人机交互结构的问题,引入HRIBench基准,它以结构化场景脚本定义协作任务与交互角色,含多项可解释指标。通过实验表明该基准能暴露机器人策略局限,提升协作性能,推动以交互为中心的机器人学习。

AI 中文摘要

当前的视觉-语言-动作(VLA)基准主要评估孤立的操作技能,而很大程度上未对人机交互结构进行建模。现实世界的协作根本上需要在共享代理下进行协调,包括意图理解、时间同步、协议遵守和动态环境中的安全交互。为填补这一空白,我们引入了HRIBench,一个基于可执行交互场景的意图感知人机协作诊断基准。HRIBench将协作任务表示为结构化场景脚本,明确建模代理角色、时间依赖、协调约束和人类行为分布。在此抽象基础上,定义了三个代表性交互角色。该基准包含13个角色条件任务及超650个评估情节。除了二元任务成功,还引入了以交互为中心的可解释指标。评估结果表明当前基础机器人策略在协作设置中存在不足,在HRIBench上微调可提高协作性能,在实际应用研究中也证明了该基准对推进以交互为中心的机器人学习的价值。

英文摘要

Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largely unmodeled. However, real-world collaboration fundamentally requires coordination under shared agency, including intent understanding, temporal synchronization, protocol adherence, and safe interaction in dynamic environments. To address this gap, we introduce HRIBench, a diagnostic benchmark for intent-aware human-robot collaboration based on executable interaction scenarios. HRIBench represents collaborative tasks as structured scenario scripts that explicitly model agent roles, temporal dependencies, coordination constraints, and human behavior distributions. Building on this abstraction, HRIBench defines three representative interaction roles: Instructor, Collaborator, and Intruder, covering intent communication, joint coordination, and robustness under human intervention. The benchmark contains 13 role-conditioned tasks with over 650 evaluation episodes generated from diverse interaction trajectories and scene variations. Beyond binary task success, HRIBench introduces interpretable interaction-centric metrics spanning synchronization, responsiveness, protocol compliance, and safety. We evaluate adapted policies based on GR00T, pi0.5, and ACT under a unified protocol. Results show that current foundation robot policies struggle substantially in collaborative settings despite strong manipulation ability, revealing major limitations in temporal coordination and intent-aware behavior. Fine-tuning on HRIBench consistently improves collaborative performance. In a real-world adaptation study, simulation data generated by HRIBench improves GR00T N1.5's physical-task success rate from 0.10 to 0.43, demonstrating the benchmark's value for advancing interaction-centric robot learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑