面向真实环境感知的零样本文本到语音合成:基于解耦音频填充的研究
Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling
浏览论文内容
中文总结 AI 辅助
本文提出扩展版DAIEN-TTS框架,解耦建模语音、噪声与混响,通过分离模块与引导机制实现环境感知零样本TTS,实验验证其自然度、相似度与可控性优于现有系统。
中文摘要 AI 辅助
当前的零样本文本到语音(TTS)系统已实现出色的自然度和说话人相似度,但通常需要高质量的说话人提示,且要么剥离声学环境信息,要么将其与说话人特征纠缠在一起,限制了其在真实场景中的适用性。本文提出了扩展版DAIEN-TTS,这是一种环境感知的零样本TTS框架,可解耦并联合建模语音、背景噪声和混响,通过独立的说话人提示和环境提示实现音色与声学环境的独立控制。该框架基于基于流匹配的F5-TTS构建,采用语音-环境分离模块将环境语音分解为语音、噪声和混响分量,再将这些分量注入扩散Transformer以实现环境感知生成。训练时使用通过将纯净语音与噪声及房间冲激响应混合构建的模拟数据,同时采用跨说话人条件策略抑制环境分支中的说话人信息泄露。当获得真实数据时,该系统可进一步微调以弥合模拟到真实的领域差距。推理阶段,一种三分支无分类器引导机制实现对语音、噪声和混响的细粒度控制,信噪比适配策略则使合成语音与环境提示对齐。在模拟和真实测试集上的实验表明,DAIEN-TTS生成的环境个性化语音具有高自然度、强说话人相似度以及忠实的噪声和混响复现能力,同时提供了超越现有环境感知TTS系统的可控性。
英文摘要
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or entangle the acoustic environment with speaker characteristics, limiting their real-world applicability. We present an extended DAIEN-TTS, an environment-aware zero-shot TTS framework that disentangles and jointly models speech, background noise, and reverberation, enabling independent control over timbre and acoustic environment through separate speaker and environment prompts. Built upon the flow-matching-based F5-TTS, it uses a speech-environment separation module to decompose environmental speech into speech, noise, and reverberation components, which are injected into the Diffusion Transformer for environment-aware generation. Training uses simulated data constructed by mixing clean speech with noise and room impulse responses, together with a cross-speaker conditioning strategy that suppresses speaker information leakage from the environment branch. When real-world data are available, the system can be further fine-tuned to bridge the simulated-to-real domain gap.At inference, a triple classifier-free guidance mechanism enables fine-grained control over speech, noise, and reverberation, and a signal-to-noise-ratio adaptation strategy aligns the synthesized speech with the environment prompt. Experiments on simulated and real-world test sets show that DAIEN-TTS generates environmental personalized speech with high naturalness, strong speaker similarity, and faithful noise and reverberation reproduction, while offering controllability beyond prior environment-aware TTS systems.