arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16514eess.AS

语音生成中的演化瓶颈:从CosyVoice到Qwen-Audio-3.0-TTS的接口协同设计与分阶段对齐

The Evolving Bottleneck in Speech Generation: Interface Co-design and Staged Alignment from CosyVoice to Qwen-Audio-3.0-TTS

  • Alibaba Token Foundry(阿里通义实验室)

机构由 AI 辅助整理,请以论文原文为准。

Qian Chen, Xiangang Li, Xiang Lv, Han Zhao, Tianyu Zhao

AI总结:

本文回顾CosyVoice系列到Qwen-Audio-3.0-TTS的演进,揭示语音生成系统通过反复迁移主导瓶颈实现进步,提出接口协同设计与分阶段对齐方法,并总结出诊断和训练模块化语音生成器的实用原则。

AI中文摘要:

语音合成系统通常被描述为一系列更大模型、更好分词器和更广泛数据的组合。这篇技术回顾性文章对CosyVoice系列(从CosyVoice到CosyVoice 2、CosyVoice 3,再到Qwen-Audio-3.0-TTS)提供了不同的解读:进展源于系统主导瓶颈的反复迁移。在整个系列中,一个稳定的分解将用于规划语音的自回归语言模型与用于渲染声学的流匹配模型分离开来。变化的是它们之间的契约。CosyVoice建立了受监督的语义分词器作为内容对齐的接口;CosyVoice 2使该接口可因果地用于流式处理,并从语言模型中移除了话语级说话人嵌入;CosyVoice 3通过多任务监督、规模扩展和可微奖励优化提高了接口的可学习性和覆盖范围;Qwen-Audio-3.0-TTS降低了分词率,使其渲染器基于连续的语言模型隐藏状态而非分词嵌入,并逐步对齐耦合系统。我们通过四个接口维度——表示、所有权、可用性和梯度可达性——将这一历史形式化,并将论文内部证据与跨论文比较区分开来。由此得到的综合框架连接了离散自回归、连续非自回归、混合和连续自回归语音生成范式,并为诊断和训练模块化语音生成器提供了实用原则。

英文摘要:

Speech synthesis systems are commonly narrated as a sequence of larger models, better tokenizers, and broader data. This technical retrospective offers a different account of the CosyVoice lineage, from CosyVoice through CosyVoice 2 and CosyVoice 3 to Qwen-Audio-3.0-TTS: progress came from repeatedly relocating the system's dominant bottleneck. Across the lineage, a stable decomposition separates an autoregressive language model that plans speech from a flow-matching model that renders acoustics. What changes is the contract between them. CosyVoice establishes supervised semantic tokens as a content-aligned interface; CosyVoice 2 makes that interface causally available for streaming and removes the utterance-level speaker embedding from the language model; CosyVoice 3 improves the learnability and coverage of the interface through multitask supervision, scaling, and differentiable reward optimization; and Qwen-Audio-3.0-TTS reduces token rate, conditions its renderer on continuous language-model hidden states instead of token embeddings, and progressively aligns the coupled system. We formalize this history through four interface dimensions---representation, ownership, availability, and gradient reach---and separate within-paper evidence from cross-paper comparison. The resulting synthesis connects discrete autoregressive, continuous non-autoregressive, hybrid, and continuous autoregressive speech-generation paradigms, and yields practical principles for diagnosing and training modular speech generators.

补充信息

↑