发表机构
NetEase Youdao(网易有道)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出支持14种语言的Confucius4-TTS系统,采用两阶段架构,无需音频提示转录即可实现跨语言零样本TTS,在多项基准测试中表现优异,相关资源已公开。
AI 中文摘要
近期零样本文本转语音(TTS)领域的进展大幅提升了语音质量与语音克隆保真度,但多数零样本TTS系统在推理阶段仍依赖音频提示的文本转录内容,这种依赖限制了跨语言语音克隆,因为真实场景中的参考音频常无转录文本。本技术报告提出Confucius4-TTS,这是支持14种语言的多语言零样本TTS系统,可在不依赖音频提示转录文本的情况下完成语内与跨语言参考克隆。Confucius4-TTS采用两阶段架构,包含文本成语义(T2S)与语义成声学(S2A)模块:基于大语言模型(LLM)的T2S模块使用可学习说话人编码器从自监督语音表示中提取音色特征;条件流匹配S2A模块将预测的语义令牌转换为梅尔频谱图,当有参考转录文本时,该模型还支持延续克隆。Confucius4-TTS在大规模多语言语音数据上训练,在公开基准测试中实现了高可懂度与说话人相似度:在CV3-Eval跨语言基准上,六个方向的平均词错误率(WER)为3.73%;在内部跨语言数据集的人工评估中,其在近期开源与商业系统中取得了最佳平均总排名,相关代码、模型检查点与演示已在指定网址发布。
英文摘要
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time. This dependency limits cross-lingual voice cloning, since in-the-wild reference audio is often untranscribed. In this technical report, we present Confucius4-TTS, a multilingual zero-shot TTS system that supports 14 languages and performs both intra-lingual and cross-lingual reference cloning without requiring transcripts of audio prompts. Confucius4-TTS follows a two-stage architecture, consisting of text-to-semantic (T2S) and semantic-to-acoustic (S2A) modules. The LLM-based T2S module uses a learnable speaker encoder to extract timbre features from self-supervised speech representations, and the conditional flow-matching S2A module converts the predicted semantic tokens into mel-spectrograms. The same model also supports continuation cloning when a reference transcript is available. Confucius4-TTS is trained on large-scale multilingual speech data. It achieves high intelligibility and speaker similarity on public benchmarks. On the CV3-Eval cross-lingual benchmark, Confucius4-TTS obtains an average WER of 3.73% across six directions. On our internal cross-lingual set, it achieves the best average overall rank in human evaluation among recent open-source and commercial systems. We release code, model checkpoints, and demos at https://github.com/netease-youdao/Confucius4-TTS.
Comments12 pages, 1 figure, 6 tables