AuEmoChat:用于对话语音合成的真实情感理解与呈现
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
浏览论文内容
中文总结 AI 辅助
针对对话语音合成中情感表达不真实及多模态令牌干扰问题,提出AuEmoChat框架,通过AuEmoCodec学习情感令牌空间、AuEmoToMe合并冗余令牌,集成到模型并结合情感流匹配渲染语音,实验证明其性能优于现有基线。
中文摘要 AI 辅助
对话语音合成(CSS)旨在在用户与智能体交互中合成具有类人情感表达和上下文一致性的语音。现有CSS方法因预定义情感标签空间有限,难以呈现真实人类情感,多轮对话历史中的冗余多模态令牌也干扰上下文理解。为此提出AuEmoChat框架,开发AuEmoCodec从大规模情感语音中学习离散真实情感令牌空间,还提出AuEmoToMe合并冗余令牌,集成到自回归文本语音模型预测目标情感和语音令牌,最后通过情感流匹配渲染语音。在NCSSD - EmCap数据集上的实验表明,AuEmoChat优于现有基线,能生成更具表现力和真实感的情感语音。
英文摘要
Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predefined emotion label spaces (e.g., seven emotion categories), while redundant multimodal tokens in multi-turn dialogue history interfere with context understanding. To address these issues, we propose AuEmoChat, a CSS framework for authentic emotion understanding and rendering. First, we develop AuEmoCodec, which learns a discrete authentic emotion token space from large-scale emotional speech via finite scalar quantization, enabling a more authentic emotion representation than limited basic emotion categories. Furthermore, we propose AuEmoToMe, an authentic-emotion-guided token merging algorithm that merges redundant tokens in multimodal dialogue history while preserving emotion-relevant context. We integrate it into an autoregressive text-speech model to predict the target authentic emotion token and speech tokens. Finally, we propose Authentic Emotion Flow Matching, which renders speech by jointly conditioning on merged dialogue context, target authentic emotion, and acoustic priors. Extensive experiments on the NCSSD-EmCap dataset demonstrate that AuEmoChat outperforms state-of-the-art CSS baselines and generates more expressive and authentic emotional speech. The code and speech demos will be available at: https://github.com/AI-S2-Lab/AuEmoChat.