发表机构
Graduate School of Informatics, Kyoto University; Computer Science Department, Boise State University(京都大学情报学研究科; 博伊西州立大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出AffectLoop系统,在Misty II机器人上实现多模态情感动态感知的说话人-听话人共情交互,经试点研究验证可提升共情响应与用户满意度。
AI 中文摘要
共情社交机器人不仅应响应用户所说内容,还应响应用户在交互过程中情绪的动态变化。然而,现有的共情对话系统通常以文本为中心,主要将共情建模为从用户情绪到系统响应的单向映射,限制了其捕捉具身说话人-听话人情感交换的能力。我们提出AffectLoop,一种在Misty II机器人上实现的具情感动态感知的多模态说话人-听话人口语对话系统。该系统跟踪说话人的言语和面部情感动态,估计机器人作为听话人的自身言语和行为情感状态,并基于这两个情感流来条件化基于LLM的响应生成。随后机器人生成简短的口语共情响应,同时伴随情感一致的具身行为,形成闭合的说话人-听话人情感循环。我们在一项包含5名参与者的试点被试内研究中评估该系统,将其与省略说话人和听话人情感状态输入的、仅以话语为条件的相同基线系统进行比较。所提出的系统获得了更高的整体印象评分,尤其是在共情响应和用户满意度方面。事后日志分析进一步显示,该系统实现了更高的说话人-听话人情感对齐和更强的基于效价的痛苦恢复。这些初步结果表明,显式建模说话人情感动态和听话人情感状态可改善具身共情交互。
英文摘要
Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically evolve during interaction. However, existing empathetic dialogue systems are often text-centered and primarily model empathy as a one-way mapping from the user's emotion to the system response, limiting their ability to capture embodied speaker--listener affective exchange. We present AffectLoop, a multimodal speaker-listener emotion-dynamics-aware spoken dialogue system implemented on the Misty II robot. The system tracks the speaker's verbal and facial affective dynamics, estimates the robot listener's own verbal and behavioral affective state, and conditions LLM-based response generation on both affective streams. The robot then generates a short spoken empathetic response together with emotionally congruent embodied behavior, forming a closed speaker--listener affective loop. We evaluate the system in a pilot within-subject study with five participants, comparing it with an otherwise identical utterance-conditioned baseline that omits the speaker- and listener-affective-state inputs. The proposed system received higher overall impression ratings, especially for empathetic response and user satisfaction. Post-hoc log analysis further showed higher speaker-listener affective alignment and stronger valence-based distress recovery. These preliminary results suggest that explicitly modeling both speaker emotional dynamics and listener affective state can improve embodied empathetic interaction.
CommentsThis paper has been accepted for presentation at APSIPA ASC 2026