arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30325cs.CLcs.SD

序列轨迹与同时融合:用于指令跟随式语音合成的多情感建模

Sequential Trajectories and Simultaneous Blending: Multi-Emotion Modeling for Instruction-Following TTS

Yan Zhou, Yun Hong, Yang Feng

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对指令跟随式TTS的多情感控制问题,提出HybridEmo后训练框架,结合分组相对策略优化与混合奖励,在MultiEmo-Test上提升了轨迹正确性与融合强度,表现优于多数对比模型。

中文摘要 AI 辅助

自然语言指令可实现对合成语音的灵活控制,但情感语音合成(TTS)系统主要针对单句级情感进行建模,多情感控制的研究尚不够充分。本文研究两项互补的多情感TTS任务:一是情感轨迹,即覆盖多个有序情感阶段;二是情感融合,即多种情感在整句中共存。这些任务暴露出监督不匹配问题:监督微调(SFT)未明确评估情感特征,而单情感奖励既无法为轨迹完成提供结构感知反馈,也无法为融合提供成对感知反馈。本文提出HybridEmo,这是一种后训练框架,先用SFT初始化两项任务,再通过带样本感知的混合奖励,利用分组相对策略优化(Group Relative Policy Optimization)对齐语音token策略。对于轨迹样本,段对齐一致性结合平均证据与最弱阶段证据,以保留指定阶段的正确性与完整性;对于融合样本,基于高斯混合模型(GMM)的奖励结合离线情感空间中目标情感锚点集合的帧级支持,以及句级弱目标边际。两个分支共享自动语音识别(ASR)奖励,并在统一策略中路由。在MultiEmo-Test数据集上,HybridEmo显著提升轨迹正确性与融合强度,且说话人相似度无明显下降;人工评估显示,相比CosyVoice 3与EmoVoice-0.5B,更偏好HybridEmo,与Qwen3-TTS的偏好接近平衡。

英文摘要

Natural-language instructions enable flexible control of synthesized speech, yet emotional TTS systems primarily model a single utterance-level affect, leaving multi-emotion control underexplored. We study two complementary multi-emotion TTS tasks: emotion trajectory, which spans several ordered affective stages, and emotion blending, in which multiple emotions coexist throughout an utterance. These tasks expose a supervision mismatch: supervised fine-tuning (SFT) does not explicitly evaluate emotion features, while single-emotion rewards provide neither structure-aware feedback for trajectory completion nor pair-aware feedback for blending. We introduce HybridEmo, a post-training framework that initializes both tasks with SFT and then aligns the speech-token policy through Group Relative Policy Optimization using a sample-aware hybrid reward. For trajectory samples, segment-aligned consistency combines average and weakest-stage evidence to preserve the correctness and completeness of prescribed stages. For blending samples, a GMM-based reward combines frame-level support from the union of target-emotion anchors in an offline emotion space with an utterance-level weaker-target margin. Both branches share an ASR reward and are routed within a unified policy. On MultiEmo-Test, HybridEmo significantly improves trajectory correctness and blending intensity, without a noticeable degradation in speaker similarity. Human evaluation prefers HybridEmo to CosyVoice 3 and EmoVoice-0.5B, with nearly balanced preferences against Qwen3-TTS.

发表机构

  • Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS)(中国科学院计算技术研究所智能信息处理重点实验室)
  • State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所人工智能安全国家重点实验室)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑