发表机构
Nankai University(南开大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有TTS系统难以实现单句内细粒度情感与时长控制的问题,提出统一后训练框架,结合监督微调与强化学习,实现自然语言控制的片段级情感和时长调节,显著提升可控性并保持语音质量。
AI 中文摘要
有声书旁白、对话代理和视听配音需要能够在单个话语内传达变化的情感并调整节奏的语音。然而,大多数现有的文本到语音(TTS)系统通常依赖于话语级别的风格条件化,使得这种细粒度控制难以实现。鉴于这一点,并受到大型语言模型中后训练成功的启发,我们提出了一个统一的后训练框架,该框架为预训练的文本到语音模型赋予了对片段级情感和时长的自然语言控制能力。监督微调建立了指令条件化的语音生成,而基于组相对策略优化的强化学习则利用情感和时长奖励以及内容和说话人保持目标来细化控制精度。通过重用预训练架构,我们的方法避免了额外的推理时控制模块。实验表明,在保持语音可懂度和说话人身份的同时,细粒度可控性显著提高,凸显了后训练作为扩展现有语音合成模型的实用途径。
英文摘要
Audiobook narration, conversational agents, and audiovisual dubbing require speech that conveys changing emotions and adapts its pacing within a single utterance. But most existing TTS systems typically rely on utterance-level style conditioning, making such fine-grained control difficult to achieve. In light of this, and inspired by the success of post-training in large language models, we propose a unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration. Supervised fine-tuning establishes instruction-conditioned speech generation, while reinforcement learning with group relative policy optimization refines control accuracy using emotion and duration rewards alongside content and speaker preservation objectives. By reusing the pretrained architecture, our approach avoids additional inference-time control modules. Experiments demonstrate significantly improved fine-grained controllability while maintaining speech intelligibility and speaker identity, highlighting post-training as a practical approach to extending existing speech synthesis models.