arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11570cs.ROcs.HC

ERR@HRI 3.0 挑战赛:人机交互中错误与预期的多模态检测

ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions

Maria Teresa Parreira, Micol Spitale, Maia Stiber, Shiye Cao, Amama Mahmood, Chien-Ming Huang, Hatice Gunes, Wendy Ju

首次发表
浏览论文内容

中文总结 AI 辅助

ERR@HRI 3.0 挑战赛提供自然场景视频数据集,供研究者开发多模态机器学习模型检测人机交互错误与预期,三个团队提交的有效模型超卷积神经网络基线,为构建相关检测系统提供了数据、任务、基线及结果等参考。

中文摘要 AI 辅助

随着机器人越来越融入人类环境,其检测和应对错误的能力对于维持用户信任和交互质量至关重要。尽管机器学习的最新进展提高了错误检测能力,但大多数方法仅限于特定上下文、受控设置或预提取特征,限制了其在现实世界条件下的通用性和适用性。为应对这一挑战,ERR@HRI 3.0 挑战赛为研究人员提供了两个互补数据集,以实现人机交互中错误检测和预防方法的端到端创新。挑战赛提供了来自自然场景的原始、非匿名视频数据:(1)旁观者影响检测(BAD)数据集,包含 45 名参与者对机器人和人类失败场景的自发反应的网络摄像头记录;(2)坏主意数据集,包含 29 名参与者在预测失败发生前的行动结果时的预期面部反应。两个数据集均通过众包收集,捕捉了现实世界条件的固有变异性。这种自然变异性虽然具有挑战性,但为开发强大的错误检测系统提供了一个真实的测试平台。参与者开发了用于旁观者反应检测(赛道 1)和预期结果预测(赛道 2)的多模态机器学习模型,以及一个可选的跨数据集泛化赛道(赛道 3)。三个团队提交了有效模型,所有模型均超过了我们的卷积神经网络基线。本文描述了 ERR@HRI 3.0 的数据集、任务、基线和结果,并讨论了对构建用于人机交互的通用、上下文感知和预期错误检测系统的意义。

英文摘要

As robots become increasingly integrated into human environments, their ability to detect and respond to errors remains critical for maintaining user trust and interaction quality. While recent advances in machine learning have improved error detection capabilities, most approaches are limited to specific contexts, controlled settings, or pre-extracted features, limiting their generalizability and applicability to real-world conditions. To address this challenge, the third edition of the ERR@HRI Challenge (ERR@HRI 3.0) provided researchers with two complementary datasets that enable end-to-end innovation in methods for both detecting and preventing errors in human-robot interaction. The challenge offered raw, non-anonymized video data from naturalistic settings: (1) the Bystander Affect Detection (BAD) dataset, containing webcam recordings of 45 participants' spontaneous reactions to robot and human failure scenarios; and (2) the Bad Idea dataset, featuring 29 participants' anticipatory facial responses while predicting action outcomes before failures occur. Both datasets were collected via crowdsourcing, capturing the inherent variability of real-world conditions. This naturalistic variability, while challenging, provides an authentic testbed for developing robust error detection systems. Participants developed multimodal machine learning models for bystander reaction detection (Track 1) and anticipatory outcome prediction (Track 2), with an optional cross-dataset generalization track (Track 3). Three teams submitted valid models, all of which surpassed our convolutional neural network baselines. This paper describes the datasets, tasks, baselines, and results of ERR@HRI 3.0, and discusses implications for building generalizable, context-aware, and anticipatory error detection systems for human-robot interaction.

发表机构

  • Cornell University(康奈尔大学)
  • Microsoft Research(微软研究院)
  • Johns Hopkins University(约翰霍普金斯大学)
  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

↑