多模态语音活动预测用于社交机器人调解:预期行为与部署约束
Multimodal Voice Activity Projection for Social Robot Mediation: Expected Behavior and Deployment Constraints
浏览论文内容
中文总结 AI 辅助
本文提出MM-VAP模型,利用视听编码器和LoRA适配预测社交机器人调解中的人际轮次动态,实验验证其可行性,并探讨部署约束。
中文摘要 AI 辅助
轮次预测对于在人际互动中充当调解者的社交机器人尤为重要,其预期行为往往不是发言,而是调整方向、等待、避免打断或准备平衡的干预。本文提出多模态语音活动预测(MM-VAP),作为未来机器人调解行为的人类状态感知层。该模型从同步的视听证据中估计会话主导权的未来演变,并推导出轮次转换事件,如保持、转移、转移预测、反馈预测及重叠相关状态。该方法采用基于语音活动相关的预训练视听编码器、LoRA适配、说话者间注意力以及从未来语音活动投影进行零样本事件推断。在NoXi、NoXi+J和Haru EDR上的实验支持了该公式的可行性,尤其是对于可与注视准备、主动倾听和保守干预相关联的主导权管理事件。最后,论文定义了预期的机器人输出接口,并讨论了主要部署约束,包括实时推断、预处理延迟、多模态同步和输入质量监控。
英文摘要
Turn-taking prediction is especially relevant for social robots that act as mediators in human-human interaction, where the expected action is often not to speak, but to orient, wait, avoid interruption, or prepare a balanced intervention. This paper presents Multimodal Voice Activity Projection (MM-VAP) as a human state-aware perception layer for future robot mediation behavior. The model estimates the future evolution of the conversational floor from synchronized audio-visual evidence and derives turn-taking events such as Hold, Shift, Shift prediction, Backchannel prediction, and overlap-related states. The approach uses VA-related pretrained audio-visual encoders, LoRA adaptation, inter-speaker attention, and zero-shot event inference from future voice activity projections. Experiments on NoXi, NoXi+J, and Haru EDR support the feasibility of this formulation, especially for floor management events that can be connected to gaze preparation, active listening, and conservative intervention. Finally, the paper defines the expected robot output interface and discusses the main deployment constraints, including real-time inference, preprocessing latency, multimodal synchronization, and input-quality monitoring.
发表机构
- i Intelligent Insights(4i智能洞察)
- Universidad de Sevilla(塞维利亚大学)
- Universidad Pablo de Olavide(巴勃罗·德·奥拉维德大学)
- Honda Research Institute Japan(本田研究院日本)
机构由 AI 辅助整理,请以论文原文为准。