突破15%的障碍:一个由用户非语言线索触发的、用于主动社交机器人的真实世界数据驱动系统
Breaking the 15% Barrier: A Real-World Data-Driven System for Proactive Social Robot Triggered by User Nonverbal Cues
浏览论文内容
中文总结 AI 辅助
研究零售商店中服务机器人交互,发现非语言行为触发机器人话语占比15.3%。通过分析客户行为定义非语言线索,开发识别器和对话框架,能基于非语言线索实现主动机器人响应,还进行了离线评估和展示在线原型。
中文摘要 AI 辅助
零售商店中的服务机器人越来越依赖级联语音管道(STT-LLM-TTS),但许多客户与机器人的交互是由接近、挥手、指向或展示物品等非语言行为发起或引导的。本文通过远程操作的人形机器人在真实商店部署中研究了这些线索,发现机器人的相当一部分轮次是由非语言行为而非语音输入触发的,揭示了仅音频对话系统的局限性。在为期6天的实地部署中,15.3%的机器人话语是由用户的非语言行为而非语音输入发起的。基于对观察到的客户行为的分析,定义了一组频繁的、与服务相关的非语言线索,并开发了一个从视频在线运行的实时多人、多标签识别器。然后提出了一个对话框架,该框架基于识别出的非语言线索令牌来调整基于大语言模型的话语生成,并在展示物品时可选地利用视觉语言模型,实现无需手工规则的主动机器人响应。离线评估了该方法在非语言触发轮次上的效果,并展示了一个实时响应用户非语言线索的在线原型。
英文摘要
Service robots in retail stores increasingly rely on cascaded speech pipelines (STT-LLM-TTS), yet many customer-robot interactions are initiated or guided by nonverbal behaviors such as approaching, waving, pointing, or showing items. This paper studies such cues in a real-world store deployment with a teleoperated humanoid robot and shows that a non-negligible portion of robot turns are triggered by nonverbal behaviors rather than spoken input, revealing a limitation of audio-only dialogue systems. In a 6-day in-the-wild deployment, 15.3\% of robot utterances were initiated by users' nonverbal behaviors rather than spoken input. Based on an analysis of observed customer behaviors, we define a set of frequent, service-relevant nonverbal cues and develop a real-time multi-person, multi-label recognizer that runs online from video. We then propose a dialogue framework that conditions LLM-based utterance generation on recognized nonverbal cue tokens, and optionally leverages a vision-language model when items are shown, enabling proactive robot responses without hand-crafted rules. We evaluate the approach offline on nonverbal-triggered turns and demonstrate an online prototype that reacts to users' nonverbal cues in real time.
发表机构
- Kyushu Institute of Technology(九州工业大学)
- CyberAgent(CyberAgent公司)
- The University of Osaka(大阪大学)
机构由 AI 辅助整理,请以论文原文为准。