发表机构
Shanghai Jiao Tong University; Shanghai Innovation Institute; Tencent; Zhejiang University(上海交通大学; 上海创新研究院; 腾讯; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出GROW方法,将其应用于DiTAR模型,在LibriSpeech和Seed-TTS数据集上提升TTS的WER和说话人相似度,且训练速度更快,计划开源相关代码与模型。
AI 中文摘要
针对流匹配文本到语音(TTS)的强化学习因确定性常微分方程(ODE)采样而变得复杂:轨迹级策略梯度方法通常会将ODE转换为随机微分方程(SDE)并跟踪每步似然比,这会引入随机扰动并产生大量开销。我们提出GROW,一种直接作用于标准流匹配目标的组相对优势加权在线策略强化学习方法。对于每个提示,GROW采样一组在线策略语音,在组内分别标准化可懂度和说话人相似度奖励,并将它们结合起来重新加权流匹配回归。Wasserstein-2速度惩罚将更新后的模型锚定到冻结的预训练参考模型。引入组平均奖励基线以将奖励加权转换为优势加权。对于奖励集中的强预训练TTS模型,正指数加权由与奖励无关的自模仿主导,而零均值符号优势则保留了有效的组内信用分配。在DiTAR上实例化并在LibriSpeech和Seed-TTS英/中文数据集上评估,GROW将平均词错误率(WER)从2.016降至1.558,将说话人相似度从0.676提升至0.715,同时保持通用文本到语音质量评估(UTMOS)指标不变。使用10-NFE训练迭代和32-NFE评估时,GROW在训练速度比32-NFE DiTAR-GRPO快2.9倍的同时保持可比性能。我们将开源完整的GROW代码、忠实的DiTAR复现版本以及所有模型检查点。
英文摘要
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturbations and substantial overhead. We propose GROW, a group-relative advantage-weighted on-policy RL method that acts directly on the standard flow-matching objective. For each prompt, GROW samples a group of on-policy utterances, separately standardizes intelligibility and speaker-similarity rewards within the group, and combines them to reweight flow-matching regression. A Wasserstein-2 velocity penalty anchors the updated model to a frozen pretrained reference. A group-mean reward baseline is introduced to convert reward weighting into advantage weighting. For strong pretrained TTS models with concentrated rewards, positive exponential weighting is dominated by reward-agnostic self-imitation, whereas a zero-mean signed advantage preserves effective within-group credit assignment. Instantiated on DiTAR and evaluated on LibriSpeech and Seed-TTS EN/ZH, GROW reduces average WER from 2.016 to 1.558 and raises speaker similarity from 0.676 to 0.715 while keeping UTMOS. With 10-NFE training rollouts and 32-NFE evaluation, GROW retains comparable performance while training 2.9x faster than 32-NFE DiTAR-GRPO. We will open-source complete GROW codes, faithful DiTAR reproduction, and all model checkpoints.