AI 中文总结
针对大型音频语言模型缺乏情绪调节机制的问题,提出心理学基础的ER-EDF框架,解耦情绪感知与调节,在五个模型和两个数据集上显著提升共情对话质量。
AI 中文摘要
口语对话系统中的共情反应生成需要同时具备准确的情绪感知和适当的情绪调节。基于知觉-行动模型和情绪调节理论等心理学理论,有效的共情不仅依赖于推断用户的情绪状态,还依赖于调节该状态在反应中的表达方式。然而,近期的大型音频语言模型(LALMs)大多将情绪视为直接的条件信号,缺乏明确的调节机制,这常常导致情绪镜像而非校准支持。我们提出了ER-EDF,一个以心理学为基础的框架,在LALMs中明确解耦情绪感知和情绪调节。感知模块跟踪用户的情绪状态,而调节模块决定该状态应如何指导共情反应的生成。该框架是模型无关的,可以无缝集成到现有的LALMs中。我们进一步构建了一个口语共情对话数据集,并引入了超越词汇匹配的共情感知评估指标。在五个LALMs和两个数据集上的实验表明,ER-EDF在自动评估和人工评估中均持续提高了共情反应质量,凸显了在口语共情对话系统中联合建模情绪感知和调节的重要性,为基于心理学的共情AI开辟了新方向。
英文摘要
Empathetic response generation in spoken dialogue systems requires both accurate emotion perception and appropriate emotion regulation. Grounded in psychological theories such as the Perception-Action Model and emotion regulation theory, effective empathy depends not only on inferring a user's affective state but also on regulating how it is expressed in responses. However, recent large audio-language models (LALMs) largely treat emotion as a direct conditioning signal, lacking explicit regulatory mechanisms, which often leads to affect mirroring rather than calibrated support. We propose ER-EDF, a psychology-grounded framework that explicitly decouples emotion perception and emotion regulation in LALMs. Perception tracks the user's emotional state, while regulation determines how this state should guide empathetic response generation. The framework is model-agnostic and integrates seamlessly into existing LALMs. We further construct a spoken empathetic dialogue dataset and introduce empathy-aware evaluation metrics beyond lexical matching. Experiments across five LALMs and two datasets show that ER-EDF consistently improves empathetic response quality in both automatic and human evaluations, highlighting the importance of jointly modeling emotion perception and regulation in spoken empathetic dialogue systems, paving a new direction for psychologically grounded empathetic AI.