FireRedTTS3:基于语义丰富的语音表征的统一语音生成与编辑
FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations
- Xiaohongshu(小红书)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出FireRedTTS3框架,利用冻结音频编码器正则化特征空间缓解自回归误差,其两个变体在对应任务数据集上均优于竞争系统,实现稳定可控的高保真语音生成与编辑。
AI中文摘要:
近期的连续自回归TTS模型直接在连续语音表征上运行,在保留丰富声学细节的同时利用了文本大语言模型的指令遵循能力。这一范式为语音克隆、指令控制的语音设计和语音编辑开辟了新可能,但仍易在自回归生成过程中出现误差累积。现有解决方案通常需要额外的语义模块、多阶段分词器训练流程或复杂的自回归架构。本研究提出了FireRedTTS3,一种简单却有效的语音生成与编辑框架,可在表征层面缓解误差累积。具体而言,我们利用在多样语音理解任务上训练的冻结音频编码器作为语义教师,对音频特征空间进行正则化,这提升了文本-语音对齐效果并稳定了自回归生成,同时保持整体系统的简洁性。FireRedTTS3提供两个变体:FireRedTTS3-Base用于多语言和多方言零样本语音克隆,FireRedTTS3-Instruct用于统一语音克隆、指令控制的语音设计和语音编辑。实验表明,在Seed-TTS-Eval和MiniMax-MLS-Test数据集上,FireRedTTS3-Base在对比系统中取得了最佳的平均语音可懂度和说话人相似度;而FireRedTTS3-Instruct在InstructTTSEval和Ming-Freeform-Audio-Edit数据集上的表现优于竞争系统。这些结果证明,语义丰富的连续语音表征结合简单架构,可实现稳定、可控且高保真的语音生成与编辑,代码和模型可在该https网址获取。
英文摘要:
Recent continuous autoregressive TTS models operate directly on continuous speech representations, preserving rich acoustic details while leveraging the instruction-following capabilities of text LLMs. This paradigm opens new possibilities for voice cloning, instruction-controlled voice design, and speech editing, but remains susceptible to error accumulation during autoregressive generation. Existing solutions often require additional semantic modules, multi-stage tokenizer training pipelines, or complex autoregressive architectures. In this work, we propose FireRedTTS3, a simple yet effective speech generation and editing framework that mitigates error accumulation at the representation level. Specifically, we leverage a frozen Audio Encoder trained on diverse speech understanding tasks as a semantic teacher to regularize the audio feature space. This improves text-speech alignment and stabilizes autoregressive generation while keeping the overall system simple. FireRedTTS3 provides two variants: FireRedTTS3-Base for multilingual and multi-dialect zero-shot voice cloning, and FireRedTTS3-Instruct for unified voice cloning, instruction-controlled voice design, and speech editing. Experiments show that FireRedTTS3-Base achieves the best average speech intelligibility and speaker similarity among compared systems on Seed-TTS-Eval and MiniMax-MLS-Test, while FireRedTTS3-Instruct outperforms competing systems on InstructTTSEval and Ming-Freeform-Audio-Edit. These results demonstrate that semantically enriched continuous speech representations, combined with a simple architecture, enable stable, controllable, and high-fidelity speech generation and editing. Code and models are available at https://github.com/FireRedTeam/FireRedTTS3.