arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36324cs.SDcs.LG

局部蒸馏,全局调度:用于少步文本到语音的流映射

Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech

  • Institute of Foundation Models (IFM), MBZUAI(基础模型研究所(IFM),穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

Yentl Collin, Evan Dufraisse, Amr Mohamed, Amine Khelif Khelif, Dani Bouch, Guokan Shang

AI总结:

本文提出局部流映射蒸馏(LFMD)及TD-DP采样调度,避免教师轨迹积分,在少步TTS中实现低NFE合成,于Seed-TTS上1-NFE达1.80% WER,接近32-NFE教师水平。

AI中文摘要:

流匹配文本到语音(TTS)模型实现了高合成质量,但需要大量神经函数评估(NFE)来积分其生成轨迹。近期针对TTS的少步流映射蒸馏方法从数值积分的教师轨迹构建目标,在目标准确性与训练成本之间产生权衡。我们提出局部流映射蒸馏(LFMD),它将欧拉映射蒸馏适配到条件TTS,并在目标构建过程中避免教师轨迹积分。对于推理,我们从教师动力学和学习映射的一致性中推导出采样调度(TD-DP),单个成本图支持多种NFE预算,无需外部音频度量评估。由于调度在单次NFE下不提供灵活性,我们通过使用软DTW的对齐感知时间自蒸馏来优化该机制。在Seed-TTS和LibriSpeech-PC上,LFMD相比匹配的积分蒸馏基线改善了低NFE合成。在Seed-TTS上,精炼学生模型在1-NFE下达到1.80%的词错误率(WER),而其32-NFE教师模型为1.76%。

英文摘要:

Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their generative trajectories. Recent few-step flow-map distillation approaches for TTS construct targets from numerically integrated teacher trajectories, creating a trade-off between target accuracy and training cost. We propose Local Flow-Map Distillation (LFMD), which adapts Eulerian Map Distillation to conditional TTS and avoids teacher trajectory integration during target construction. For inference, we derive a sampling schedule (TD-DP) from teacher dynamics and consistency of the learned maps, with a single cost graph supporting multiple NFE budgets without external audio-metric evaluation. Because scheduling offers no flexibility at one NFE, we refine this regime with alignment-aware temporal self-distillation using soft-DTW. Across Seed-TTS and LibriSpeech-PC, LFMD improves low-NFE synthesis over a matched integral-distillation baseline. On Seed-TTS, the refined student reaches 1.80% WER with 1-NFE, compared with 1.76% for its 32-NFE teacher.

补充信息

↑