发表机构
Boise State University; Semio; Peerbots; Georgia Institute of Technology(博伊西州立大学; Semio公司; Peerbots公司; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一个可扩展的社交机器人语音对话基准,涵盖澄清、打断、具身信号和时间约束,并采用实时通信框架Retico以支持人机语音交互。
AI 中文摘要
语言模型为人类与机器人之间提供了即插即用的接口,但当需要语音、对话、快速交互和协作时,仍存在重要挑战。我们提出了一个供社区使用的基准测试,用以探索机器人与人类之间常见的语音对话产物,包括澄清请求、打断、具身信号(如点头或面部线索)以及时间约束。我们还阐述了将基准测试扩展到其他对人机交互重要方面的愿景,这些方面对更广泛的研究社区具有重要意义。为促进该基准测试,我们进一步提出使用\textit{Retico},一个实时通信框架,它满足了使机器人具备语音对话能力的重要技术要求。
英文摘要
Language models provide a plug-and-play interface between humans and robots, but important challenges remain when speech, dialogue, fast interaction, and collaboration are required. We propose a benchmark for the community to use as a way to explore common spoken dialogue artifacts between robots and humans, including requests for clarification, interruptions, embodied signals (e.g., head nods or facial cues), and time constraints. We also explain our vision to extend the benchmark for other aspects of human-robot interaction that are important to the larger research community. To facilitate the benchmark, we further propose using \textit{Retico}, a real-time communication framework that fulfills important technical requirements to enable robots to have spoken dialogue capabilities.
CommentsAccepted to and presented at the IROS 2026 Workshop on Human-Robot Dialogue