arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EMODE:用于情感感知语音语言建模的动态副语义专家

EMODE: Dynamic Para-Semantic Experts for Emotion-Aware Speech Language Modeling

Jianan Pan, Yiwen Gu, Xinze Li, Rui Wang, Kejie Huang

arXiv 2610.06956首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EMODE提出动态副语义专家(DPSE)分解语音特征,通过三阶段课程训练,在SER和共情任务上平衡词汇保真与情感敏感,增强跨语料情感理解。

AI 中文摘要

大型语音语言模型在统一的跨模态理解与生成方面展现出强大能力,然而副语言线索,尤其是情感,仍然难以保留。现有系统通常依赖纠缠的声学表示,这使得底层语言模型过度依赖恢复的词汇内容,而非基于声学韵律证据来引导其行为。我们通过EMODE来解决这一局限,这是一种围绕动态副语义专家(DPSE)构建的情感感知语音语言模型。DPSE将连续语音特征分解为语义和副语言通路,动态路由它们,并在集成到语言模型之前进行融合。为了将这种结构分解转化为功能特化,EMODE采用三阶段课程训练,包括语义预热、副语言激活和联合精炼,并由正交专家引导(OEG)、语义-声学对齐(SAA)和门控多样性正则化(GDR)指导。在SER测试、共情反应评估以及新构建的双语MEPA基准上的实验表明,EMODE改善了词汇保真度与情感敏感性之间的平衡,增强了基于情感的响应生成,并揭示了显式副语义分解对于稳健的跨语料库情感理解的价值。

英文摘要

Large speech language models have demonstrated strong capabilities in unified cross-modal understanding and generation, yet paralinguistic cues, especially emotion, remain difficult to preserve. Existing systems typically rely on entangled acoustic representations, which allow the underlying language model to depend excessively on recovered lexical content instead of grounding its behavior in acoustic-prosodic evidence. We address this limitation with EMODE, an emotion-aware speech language model built around \textbf{Dynamic Para-Semantic Experts (DPSE)}. DPSE decomposes continuous speech features into semantic and paralinguistic pathways, routes them dynamically, and fuses them before integration into the language model. To turn this structural decomposition into functional specialization, EMODE is trained with a three-stage curriculum consisting of semantic warm-up, paralinguistic activation, and joint refinement, guided by Orthogonal Expert Guidance (OEG), Semantic-to-Acoustic Alignment (SAA), and Gating Diversity Regularization (GDR). Experiments on SER test, empathetic response evaluation, and the newly constructed bilingual MEPA benchmark show that EMODE improves the balance between lexical fidelity and emotional sensitivity, strengthens affect-grounded response generation, and exposes the value of explicit para-semantic factorization for robust cross-corpus emotion understanding.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑