arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

嘈杂公共空间中多参与者的人机对话

Human-robot conversation with multiple participants in noisy public spaces

Divesh Lala, Yogeeswaran Muthukumaran, Vincent Fernandes, Kazushi Kato, Shota Fujiki, Zihao Chi, Masaya Iwasaki, Taiken Shintani, Megumi Kawata, Kazuki Sakai, Koji Inoue, Yuicihiro Yoshikawa, Tatsuya Kawahara

arXiv 2609.00648首次发表:更新:

发表机构

Kyoto University; Osaka University; National University of Singapore; ENSEA(京都大学; 大阪大学; 新加坡国立大学; 法国国立高等先进技术学校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出适用于嘈杂公共空间多参与者人机对话的音频系统,可用于人形机器人ERICA的专注聆听及移动Teleco机器人的虚拟形象对话支持场景,采用单多通道麦克风阵列,增强语音并提供空间音频以提升沉浸感,已在2025大阪世博会验证。

AI 中文摘要

针对开放公共空间等嘈杂真实环境,自主机器人和虚拟形象(avatar)的口语对话系统需精心设计以提供增强语音信号,这类信号可用于语音识别,或在虚拟形象系统中作为纯净语音传输给远程操作员。本研究提出一种可适用于上述两种场景的音频系统,作为概念验证在2025年大阪世博会展示。该系统的两种应用场景为:一是与人形机器人ERICA配合的专注聆听系统,二是与移动Teleco机器人配合的对话支持系统,其中一台Teleco机器人作为远程操作员的虚拟形象,两种场景均支持多方对话,且采用单个多通道麦克风阵列。我们描述该音频系统不仅可增强嘈杂环境中多位说话人的语音,还能提供空间音频形式,提升基于虚拟形象的对话交互的沉浸感。

英文摘要

For noisy real-world environments such as those in open public spaces, spoken dialogue systems for both autonomous robots and avatars should be carefully designed to provide enhanced speech signals. These signals can be used either for speech recognition or, in the case of an avatar system, transmitted as clean speech to a remote operator. This work proposes an audio system that can be used for both these scenarios and was demonstrated as a proof-of-concept at the 2025 World Expo in Osaka. The first scenario is an attentive listening system with the android ERICA, and the second is a conversation support system with mobile Teleco robots, with one of them acting as an avatar for a remote operator. Both systems feature multi-party conversation and use a single multi-channel microphone array. We describe how our audio system not only enhances the speech of multiple speakers in a noisy environment, but provides a form of spatial audio which allows for more immersiveness in avatar-based conversational interactions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑