arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26422cs.CLeess.AS

利用对话上下文丰富语音情感表征

Enriching Speech Emotion Representations with Conversational Context

  • Université Paris-Saclay, CEA, List(巴黎萨克雷大学,法国原子能委员会,List)
  • LISN, CNRS, Université Paris-Saclay(LISN,法国国家科学研究中心,巴黎萨克雷大学)

机构由 AI 辅助整理,请以论文原文为准。

Arthur Peuvot, Romaric Besançon, Gaël de Chalendar, Bianca Vieru, Ioana Vasilescu

AI总结:

本文提出ACERT模块,通过整合灵活长度的对话上下文窗口来捕捉语音交互中的情感演变,在IEMOCAP、SAFE和MELD数据集上超越现有方法,验证了情感与对话连续性的关键作用。

AI中文摘要:

检测情感对于构建能够准确且自适应地与人类交互的系统是必要的。语音情感识别(SER)已成为开发智能口语界面的重要研究焦点。然而,大多数研究在话语层面预测情感,忽略了对话上下文及其所承载的情感流动和说话者互动。在本文中,我们引入了ACERT(随时间平均的上下文情感表征),一个整合了灵活长度对话上下文窗口的模块,以更好地捕捉口语互动中的情感演变。为了评估该方法的鲁棒性,我们在涵盖多种情感表达风格和上下文的数据集上进行了实验。ACERT在IEMOCAP上优于当前最先进(SOTA)方法,在SAFE上建立了首个上下文感知基准,并在MELD上针对未加权、类别平衡指标取得了强劲结果。消融研究表明,ACERT的性能提升来自于情感和对话的连续性,而非说话者身份或声学条件。

英文摘要:

Detecting emotions is necessary for building systems that can accurately and adaptively interact with humans. Speech Emotion Recognition (SER) has become an important research focus to develop intelligent spoken interfaces. However, most studies predict emotions at the utterance level, ignoring the conversational context, along with the emotional flow and speaker interactions it carries. In this paper, we introduce ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in spoken interactions. To evaluate the robustness of this method, we conducted experiments on datasets spanning diverse emotionally expressive styles and contexts. ACERT outperforms current state-of-the-art (SOTA) approaches on IEMOCAP, establishes the first context-aware benchmark on SAFE, and obtains strong results on MELD for unweighted, class-balanced metrics. Ablation studies show that ACERT's gains come from emotional and conversational continuity, rather than from speaker identity or acoustic conditions.

补充信息

↑