arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

声调-词义冲突下的鲁棒语音情感识别:基准与框架

Robust Speech Emotion Recognition under Tone-Word Conflict: A Benchmark and Framework

Xiaojiang Peng, Dawei Huang, Yongjie Lv, Ruijie Xiong, Chunxiang Jin, Bin Li, Xiaohui Wang, Zitong Yu

arXiv 2609.04236首次发表:更新:

发表机构

Shenzhen Technology University; Ant Group; Skyworth Digital Technology Co., Ltd; Xiaopai Technology Co., Ltd; Great Bay University(深圳技术大学; 蚂蚁集团; 创维数字技术有限公司; 小派科技有限公司; 大湾区大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对声调-词义冲突场景下现有语音情感识别模型性能严重下降的问题,提出DAS框架,构建基准TWIN-SER,实验表明其在相关场景及标准设置中均优于现有方法。

AI 中文摘要

语音情感识别(SER)是人机交互的关键组成部分,受到工业界和学术界的广泛关注。然而,现有SER系统通常假设声调与词义语义一致,忽略了现实世界中存在声调-词义冲突的场景——即语音传达的情感与词语字面含义相矛盾。为弥合这一差距,本文提出TWIN-SER(声调-词义不一致语音情感识别),这是一个用于声学-语义不一致场景下系统评估的基准,研究发现现有最先进模型在这种不一致场景下性能会严重下降。为解决该问题,本文提出DAS(解耦声学-语义融合)框架,通过显式解耦声学与语义通路、选择信息丰富的高能量嵌入,并通过轻量型基于查询的注意力机制自适应融合,来缓解声调-词义冲突。具体而言,DAS包含三个关键模块:i)异构特征提取模块,分别从原始输入中捕获互补的声学与语义表示;ii)高能量嵌入选择模块,识别并保留最具判别力的嵌入;iii)Q-Former组合模块,通过交叉注意力连接两个通路,实现不一致场景下的鲁棒情感预测。大量实验表明,DAS在声调-词义冲突场景以及标准域内、零样本设置中均始终优于现有方法。本文的代码和数据集可在该https网址获取。

英文摘要

Speech emotion recognition (SER) is a crucial component of human-computer interaction, attracting extensive attention from both industry and academia. However, existing SER systems typically assume alignment between vocal tone and lexical semantics, overlooking the real-world scenarios that involve tone-word conflict-where the emotion conveyed by speech contradicts the literal meaning of the words. To bridge this gap, we introduce TWIN-SER (Tone-Word Incongruent SER), a benchmark for systematic evaluation under acoustic-semantic incongruence, and show that state-of-the-art models degrade severely under such incongruence. To address this, we propose DAS (Disentangled Acoustic-Semantic fusion), a framework that mitigates tone-word conflict by explicitly disentangling acoustic and semantic pathways, selecting informative high-energy embeddings, and adaptively fusing them via a lightweight query-based attention mechanism. Specifically, DAS comprises three crucial modules: i) a heterogeneous feature extraction module that separately captures complementary acoustic and semantic representations from raw input; ii) a high-energy embedding selection module that identifies and retains the most discriminative embeddings; and iii) a Q-Former combination module that bridges the two pathways through cross-attention, enabling robust emotion prediction under incongruent conditions. Extensive experiments demonstrate that DAS consistently outperforms existing methods in tone-word conflict scenarios, as well as in standard in-domain and zero-shot settings. Our code and datasets are available at https://github.com/24DavidHuang/FAS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑