发表机构
School of Information Science and Technology, Lanzhou University(兰州大学信息科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReaFlow-TTS通过引入话语级随机实现潜变量并施加VAD语义,实现无需目标语音的可分级属性操控,提升合成质量并保持自然度。
AI 中文摘要
在流匹配文本到语音(TTS)中,不同的语音实现在相同生成条件下可诱导不同的目标速度。使用平方误差训练的确定性速度场预测其条件均值,从而边缘化依赖于实现的变异性。同时,对此类变异性建模并不固有地提供用于属性操作的语义可解释接口。我们提出ReaFlow-TTS,一种实现条件流匹配框架,引入话语级随机实现潜变量,并利用其在生成轨迹中调节速度预测。我们进一步在实现空间上施加效价-唤醒-支配(VAD)语义,使得在推理时无需目标语音即可进行直接且分级的属性操作。实验表明,与匹配的全掩码基线相比,合成质量得到提升,且潜变量诱导的音高、能量和时序趋势在不同初始噪声样本中可重现,提供了潜变量被用作可重用实现条件的行为证据。主观评估进一步展示了跨生成上下文的分级VAD操作,且自然度仅有适度变化。
英文摘要
In flow-matching text-to-speech (TTS), different speech realizations can induce different target velocities under the same generation conditions. A deterministic velocity field trained with squared error predicts their conditional mean, thereby marginalizing realization-dependent variation. Meanwhile, modeling such variation does not inherently provide a semantically interpretable interface for attribute manipulation. We propose ReaFlow-TTS, a realization-conditioned flow-matching framework that introduces an utterance-level stochastic realization latent and uses it to condition velocity prediction throughout the generation trajectory. We further impose valence-arousal-dominance (VAD) semantics on the realization space, enabling direct and graded attribute manipulation without target speech at inference. Experiments demonstrate improved synthesis quality over a matched full-mask baseline and reproducible latent-induced pitch, energy, and timing tendencies across initial-noise samples, providing behavioral evidence that the latent is used as a reusable realization condition. Subjective evaluation further demonstrates graded VAD manipulation across generation contexts with only modest changes in naturalness.