arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11737cs.SD

Phoenix TTS:基于流匹配驱动的语音分词实现高保真合成与语音转换

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization

Peijie Chen, Zhuanling Zha, Zhipeng Nie, Weijie Wu, Yiming Liu, Daiyu Huang, Junbo Li, Jun Fang, Naiqiang Tan, Hua Chai, Qingyang Hong

首次发表
浏览论文内容

中文总结 AI 辅助

Phoenix TTS是将表示学习与流匹配驱动的生成式声学建模耦合的统一框架,经11万小时数据训练后,实现高保真语音合成,在零样本场景下兼具优异可懂度与说话人相似度,还可无缝适配零样本语音转换任务。

中文摘要 AI 辅助

在当前零样本文本转语音系统中,传统语义分词器通常通过监督自动语音识别(ASR)或自监督学习目标进行优化。然而,由于语音的固有特性,语义与声学信息无法完全解耦,基于ASR的分词器会舍弃声学细节以聚焦语言内容;依赖这些分词器的模型通常难以实现最优的说话人相似度。此外,这些分词器是独立优化的,缺乏下游声学生成任务的直接监督。这种孤立训练导致提取的离散分词与声学模型所需的连续空间之间存在特征差距,从根本上限制了合成质量的上限。为弥合这一差距,我们提出Phoenix TTS,这是一个将表示学习与生成式声学建模紧密耦合的统一框架。具体而言,我们的语音分词器被优化以重建自监督特征,从而保持语义丰富性,同时接收流匹配(Flow Matching)训练损失的直接监督。通过这种联合训练范式,提取的离散分词成功保留了必要的语义信息,并天然与下游流匹配模型的特征空间对齐。综合评估凸显了Phoenix TTS的效率与有效性:该系统在11万小时数据上训练后,实现了出色的语音可懂度,词错误率(WER)始终低于真实录音的WER;同时保持了稳健的零样本说话人相似度,可与多个知名大规模基准模型相媲美或超越。此外,作为这种统一训练的有利副产品,学习到的分词器可无缝适配零样本语音转换任务,无需特定任务的微调。

英文摘要

In current zero-shot text-to-speech systems, conventional semantic tokenizers are typically optimized using supervised automatic speech recognition or self-supervised learning objectives. However, due to the inherent nature of speech, semantic and acoustic information cannot be completely decoupled, and ASR-based tokenizers discard acoustic details to focus on linguistic content; models relying on them usually struggle to achieve optimal speaker similarity. Furthermore, these tokenizers are optimized independently and lack direct supervision from downstream acoustic generation tasks. This isolated training creates a feature gap between the extracted discrete tokens and the continuous space required by acoustic models, fundamentally bottlenecking the upper bound of synthesis quality. To bridge this gap, we propose Phoenix TTS, a unified framework that tightly couples representation learning with generative acoustic modeling. Specifically, our speech tokenizer is optimized to reconstruct self-supervised features to maintain semantic richness, while simultaneously receiving direct supervision from a Flow Matching training loss. Through this joint training paradigm, the extracted discrete tokens successfully preserve essential semantic information and natively align with the feature space of the downstream Flow Matching model. Comprehensive evaluations highlight the efficiency and effectiveness of Phoenix TTS. Trained on 110K hours of data, the system achieves excellent speech intelligibility, yielding WER that consistently falls below that of ground-truth recordings. Simultaneously, it maintains robust zero-shot speaker similarity, rivaling or outperforming several prominent large-scale baselines. Furthermore, as an advantageous byproduct of this unified training, the learned tokenizer can be seamlessly adapted to zero-shot voice conversion tasks without requiring task-specific fine-tuning.

发表机构

  • Didichuxing Co. Ltd(滴滴出行有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑