EmphTTS:一种基于强化学习的重音可控语音合成系统
EmphTTS: an emphasis-control TTS with reinforcement learning
- Aalto University(阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对语音合成中重音可控性挑战,提出EmphTTS系统,采用GRPO强化学习优化时长预测器,实现词级重音控制,显著提升重音表现与主观偏好。
AI中文摘要:
在文本到语音合成中,即使文本输入提供了明确的重音控制信号,生成可控且类人的重音仍是一个未解决的挑战,这限制了合成语音在实际应用中的交际准确性。强化学习近来在语音合成系统的后训练中展现出与人类偏好对齐的潜力,但现有方法尚未应用于词级韵律控制。我们提出EmphTTS,一种非自回归语音合成系统,将组相对策略优化(GRPO)应用于时长预测器,并配以重音定位奖励,从而实现对词级重音的直接优化。评估表明,EmphTTS在重音可控性上表现最佳,并在重音客观评估中取得最优结果。在主观偏好测试中,EmphTTS显著优于合成真实语音及大多数基线系统。消融研究显示,GRPO在重音实现上优于基于监督微调的时长建模和简单语速调整,同时缓解了独立训练的时长预测器与语音合成模型之间的不匹配问题。
英文摘要:
Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in real-world applications. Reinforcement learning has recently shown promise for post-training TTS systems to align with human preference, yet existing methods have not been applied to word-level prosodic control. We present EmphTTS, a non-autoregressive TTS system that applies Group Relative Policy Optimization (GRPO) to the duration predictor with an emphasis localization reward, enabling direct optimization for word-level emphasis. Evaluations show that EmphTTS achieves the best emphasis controllability and performs the best in emphasis objective evaluation. In subjective preference tests, EmphTTS is significantly preferred over synthetic groundtruth and most baselines. Ablation studies show that GRPO improves emphasis realization beyond supervised-finetuning-based duration modeling and simple speaking-rate adjustment, while alleviating the mismatch between the independently trained duration predictor and TTS model.