arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EchoChat:共情口语对话中的结构化认知推理

EchoChat: Structured Cognitive Reasoning in Empathetic Spoken Dialogue

Dingdong Wang, Shujie Liu, Yayue Deng, Yuxuan Hu, Yunrui Cai, Jincenzi Wu, Jianwei Yu, Jinyu Li, Helen Meng

arXiv 2610.04826首次发表:更新:

发表机构

The Chinese University of Hong Kong; Microsoft Corporation(香港中文大学; 微软公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对语音大模型在共情对话中缺乏结构化推理和错误级联问题,提出EchoChat框架,结合声学锚定注意力和步骤分解信用分配,并构建数据集与基准,实现感知、推理和反应对齐的最优性能。

AI 中文摘要

共情口语对话是一个复杂的认知过程,不仅需要识别情绪,还需要推断用户的潜在心理状态以提供适当的支持。然而,当前的语音大语言模型(SpeechLLMs)通常将共情视为直接的输入到输出的映射,导致“表面温暖”但情感空洞的交互。此外,由于共情依赖于具有强步骤间依赖性的多阶段过程,任何中间步骤的错误都可能级联到后续步骤并导致不恰当的反应,而现有的训练范式缺乏精确定位和改善此类错误的机制。在这项工作中,我们提出了EchoChat,一个统一框架,将共情口语对话重构为整合感知、心理状态推理和反应生成的结构化认知推理过程。为了支持这一范式,我们首先构建了EchoDialogue-400K,一个声学丰富的数据集,用于多阶段共情监督。在监督微调(SFT)阶段,我们通过提出的声学锚定注意力(Acoustic-Anchored Attention, AAA)来加强声学基础。在强化学习(RL)阶段,我们进一步引入了一种新颖的阶段感知优化目标,采用步骤分解信用分配(Step-Decomposed Credit Assignment, SDCA)来定位推理错误并减轻级联错误传播。此外,我们引入了EchoEval,一个专家标注的基准,用于多维度共情评估。大量实验表明,EchoChat在感知、推理和反应对齐方面达到了最先进的性能。项目页面:this https URL

英文摘要

Empathetic spoken dialogue is a sophisticated cognitive process that requires not only recognizing emotions but also inferring a user's latent mental states to provide appropriate support. However, current SpeechLLMs often treat empathy as a direct input-to-response mapping, leading to "superficially warm" but emotionally hollow interactions. In addition, since empathy relies on a multi-stage process with strong inter-step dependency, errors at any intermediate step can cascade through subsequent steps and lead to inappropriate responses, while existing training paradigms lack mechanisms to precisely localize and improve such errors. In this work, we propose EchoChat, a unified framework that reformulates empathetic spoken dialogue as a structured cognitive reasoning process integrating perception, mental-state reasoning, and response generation. To support this paradigm, we first construct EchoDialogue-400K, an acoustically rich dataset for multi-stage empathetic supervision. During the SFT stage, we strengthen acoustic grounding through proposed Acoustic-Anchored Attention (AAA). During the RL stage, we further introduce a novel stage-aware optimization objective with Step-Decomposed Credit Assignment (SDCA) to localize reasoning errors and mitigate cascaded error propagation. In addition, we introduce EchoEval, an expert-annotated benchmark for multi-dimensional empathy evaluation. Extensive experiments demonstrate that EchoChat achieves state-of-the-art performance in perception, reasoning, and response alignment. Project page: https://github.com/dingdongwang/EchoChat

CommentsNeurIPS 2026; Project page: https://github.com/dingdongwang/EchoChat

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑