AI 中文总结
JoyAI-Voice 2.0是一种全连续自回归语音生成模型,采用语义-声学双编码器联合表示,通过流匹配和强化学习优化,在词错误率和感知保真度上达到最先进水平。
AI 中文摘要
我们提出了JoyAI-Voice 2.0,一种基于全连续双编码器架构的端到端拟人语音生成模型。原始语音被编码为连续潜变量并划分为补丁。每个补丁由语义-声学双编码器分解为语义纯化表示和声学表示,两者融合后联合输入到因果自回归Transformer中进行规划。该Transformer预测下一个补丁的条件,局部扩散Transformer渲染其完整潜变量以进行48 kHz合成。模型采用联合流匹配和停止预测目标进行训练,随后进行监督微调和基于DiffusionNFT的强化学习以提升模型性能。它在Seed-TTS上实现了最低的平均词错误率2.51%,相对于最强基线相对降低了14.9%,并在InstructTTSEval上中英文属性保真度均达到最先进水平,在MDVD-Eval上10个感知维度中领先5个,总体得分最高为0.893。
英文摘要
We present JoyAI-Voice~2.0, an end-to-end anthropomorphic speech generation model built upon a fully continuous, dual-encoder architecture. Raw speech is encoded into continuous latents and partitioned into patches. Each patch is decomposed by a semantic-acoustic dual encoder into a semantically purified representation and an acoustic representation, which are fused and jointly fed to a causal autoregressive Transformer for planning. The Transformer predicts the conditioning for the next patch, and a local diffusion Transformer renders its full latents for 48\,kHz synthesis. The model is trained with a joint flow-matching and stop-prediction objective, followed by supervised fine-tuning and reinforcement learning with DiffusionNFT to improve model performance. It achieves the lowest average word error rate of 2.51\% on Seed-TTS, a 14.9\% relative reduction over the strongest baseline, and state-of-the-art attribute fidelity on InstructTTSEval in both Chinese and English, leading on 5 of 10 perceptual dimensions with the highest overall score of 0.893 on MDVD-Eval.