AI 中文总结
研究针对零样本对话语音合成的内存瓶颈问题,提出ZipL-Dialog方法,通过转移到潜在空间及采用相关技术优化,实现了大幅降低内存占用并加速推理,保持了感知自然度。
AI 中文摘要
零样本对话语音合成得益于流匹配,但在密集的梅尔频谱图上进行分钟级生成会导致严重的内存瓶颈,常常迫使进行不自然的分块合成。我们提出了ZipL-Dialog,它将条件流匹配转移到4倍时间压缩(25Hz)的潜在空间。为了在压缩下保持声学保真度,我们采用了具有辅助梅尔域监督的确定性梅尔自动编码器,并优化了ZipFormer的分层下采样调度。实验表明,ZipL-Dialog比基线将最大峰值GPU内存减少了11.22倍,并将推理速度提高了2.23倍,在保持感知自然度的同时,大幅降低了单通道多分钟对话合成的内存占用。
英文摘要
Zero-shot dialog TTS benefits from flow-matching, but minute-scale generation on dense mel-spectrograms causes severe memory bottlenecks, often forcing unnatural chunked synthesis. We propose ZipL-Dialog, which shifts conditional flow-matching into a 4x time-compressed (25 Hz) latent space. To preserve acoustic fidelity under compression, we employ a deterministic mel autoencoder with auxiliary mel-domain supervision and optimize the ZipFormer's hierarchical downsampling schedule. Experiments show that ZipL-Dialog reduces maximum peak GPU memory by 11.22x and accelerates inference by 2.23x over the baseline, substantially lowering the memory footprint of single-pass multi-minute dialog synthesis while maintaining perceptual naturalness.
CommentsAccepted to Interspeech 2026