arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

现实世界人机交互(HRI)中的多模态融洽度估计

Multimodal Rapport Estimation in Real-World HRI

Akihiro Sakuramoto, Takato Hayashi, Ryo Miyoshi, Yuki Okafuji, Shogo Okada

arXiv 2608.18401首次发表:更新:

发表机构

Japan Advanced Institute of Science and Technology; CyberAgent; The University of Osaka(日本先进科学技术学院; CyberAgent公司; 大阪大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对现实世界人机交互的融洽度估计问题,采用日本药店的62个多模态会话,对比零样本LLMs等模型,发现Gemini融合模型表现最优,且性能随互动时长、群体规模变化。

AI 中文摘要

评估现实世界人机交互(HRI)中的互动质量是一项重要挑战。若能可靠估计互动质量,结果可用于优化对话策略,最终使机器人自主调整行为。然而,现有自动评估方法主要在受控实验室环境中开发,目前尚不清楚这些方法能否直接应用于现实环境——现实环境中用户可能自由弃权(不执行),且可能自然出现多方参与的情况。本研究使用在日本一家药店收集的62个多模态记录会话,研究第三方评分的融洽度分数自动估计问题。我们对比了零样本大型语言模型(LLMs)、预训练文本、音频及视觉模型,以及它们的预测级融合。结果显示,在现实世界HRI中,零样本LLMs表现出较强性能,而音频和视觉模型往往能提供互补信息。具体而言,Gemini 2.5 Flash作为单一模型表现强劲,结合Gemini(文本)、HuBERT和V-JEPA的融合模型整体性能最佳。进一步分析表明,估计性能随互动时长和群体规模条件而变化。这些发现表明,现实世界HRI中的融洽度估计需要考虑实验室环境假设之外的情境变异性的评估与模型设计。

英文摘要

Evaluating interaction quality in real-world HRI is an important challenge. If interaction quality can be estimated reliably, the results can be used to improve dialogue strategies and ultimately enable robots to adapt their behavior autonomously. However, existing automatic evaluation methods have been developed primarily in controlled laboratory settings, and it remains unclear whether they can be directly applied to real-world environments, where users are free to disengage and multi-party participation may arise naturally. In this study, we investigate the automatic estimation of third-party-rated rapport scores using 62 sessions of multimodal recordings collected in a Japanese drugstore. We compare zero-shot LLMs, pretrained text, audio, and visual models, and their prediction-level fusion. The results show that, in real-world HRI, zero-shot LLMs achieve strong performance, while audio and visual models tend to provide complementary information. In particular, Gemini 2.5 Flash performs strongly as a single model, and a fusion model combining Gemini (text) with HuBERT and V-JEPA performs best overall. Further analyses showed that estimation performance varied across interaction-duration and group-size conditions. These findings suggest that rapport estimation in real-world HRI requires evaluation and model design that account for contextual variability beyond that assumed in laboratory settings.

Comments9 pages, 4 figures, 3 tables. Accepted at the 28th ACM International Conference on Multimodal Interaction (ICMI 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑