arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34431cs.LGcs.SD

调和条件流匹配中的频谱演化以用于文本到语音合成

Harmonizing Spectral Evolution in Conditional Flow Matching for TTS

  • Indian Institute of Technology Bombay(印度理工学院孟买分校)

机构由 AI 辅助整理,请以论文原文为准。

Isha Pandey, Varad Deshpande, Abhijat Bharadwaj, Ganesh Ramakrishnan

AI总结:

针对条件流匹配TTS推理时频谱演化不连贯的问题,提出免训练的基于离散小波变换的频率选择性增强方法,在多种架构上减少NFE并提升FAD,且不损害音质。

AI中文摘要:

用于文本到语音(TTS)的条件流匹配(CFM)模型在推理过程中存在频率演化不连贯的问题。虽然在其他领域的扩散模型中已针对类似的频谱失衡进行了处理,但这些通用解决方案无法推广到CFM固有的不协调声学动态。我们证明,通过引入一种新颖的免训练频率选择性增强策略,可以有效缓解这一问题。利用离散小波变换(DWT),我们的方法在ODE积分过程中动态调节梅尔频谱子带,通过惩罚过快的低频增长并增强滞后的高频细节来同步频谱发展。在多种架构(Matcha-TTS、F5-TTS、IndicF5)上的验证表明,我们的方法将所需的函数评估次数(NFE)从32次减少到26次,并将弗雷歇音频距离(FAD)最多提升61%,同时不影响平均意见得分、说话人相似度和语音可懂度。

英文摘要:

Conditional Flow Matching (CFM) models for text-to-speech (TTS) suffer from incoherent frequency evolution during inference. While similar spectral imbalances are addressed in diffusion models for other domains, those generic solutions fail to generalize to the inherently uncoordinated acoustic dynamics of CFM. We demonstrate that this issue can be effectively mitigated by introducing a novel training-free frequency-selective boosting strategy. Using the Discrete Wavelet Transform (DWT), our method dynamically modulates mel-spectrogram sub-bands during ODE integration, synchronizing spectral development by penalizing aggressive low-frequency growth and boosting lagging high-frequency details. Validated across diverse architectures (Matcha-TTS, F5-TTS, IndicF5), our approach reduces the required Number of Function Evaluations (NFE) from 32 to 26 and improves Frechet Audio Distance (FAD) by up to 61%, all without compromising mean opinion scores, speaker similarity, and speech intelligibility.

补充信息

↑