基于流匹配文本到语音教师的分层深度剪枝蒸馏:一个紧凑的印地语语音合成器
Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer
浏览论文内容
中文总结 AI 辅助
研究在数据量严重不足时构建紧凑印地语TTS模型的方法,通过分层深度剪枝蒸馏大型流匹配教师模型热启动学生模型,经多步剪枝与微调,得到不同参数的模型,在特定参数下效果良好,还解决了一些模型问题并进行了基准测试。
中文摘要 AI 辅助
我们提出了一种实用方法,用于在数据量严重不足(约17.6小时)的情况下,通过蒸馏一个大型流匹配教师模型(IndicF5,337M参数的DiT)来构建紧凑的印地语语音合成(TTS)模型。从头开始在这么少的数据上训练一个小模型会完全失败。相反,我们仅通过剪枝深度来从教师模型热启动学生模型:保持教师模型的宽度、文本维度、注意力头以及梅尔/文本输入输出不变,这样所有非块张量一对一复制,并保留变压器块的均匀间隔子集。我们首先测量教师模型能容忍多少深度(在-27%的块时仍接近功能正常,但超过-50%就会崩溃),然后逐渐降低(22 -> 16 -> 12 -> 8 -> 6个块),每次剪枝后重新微调,每一步都通过目标自动语音识别词错误率(WER)检查来控制。得到的学生模型在249M和190M参数时对未见句子的WER为0.00,在131M参数时仍保持稳健;在102M参数时我们观察到明显的能力悬崖,我们将其归因于数据量而非方法。我们还记录了两个训练/推理特征和库奇偶性失败(梅尔滤波器组和旋转嵌入库版本),它们会无声地降低音频质量,以及一个与版本无关的修复方法。该方法产生了一个高质量的印地语语音,可在6GB笔记本电脑GPU上实时运行。一个独立的50句FLEURS基准测试将发布的190M学生模型与其教师模型和MMS-TTS-hin进行了比较。
英文摘要
We present a practical recipe for building a compact Hindi text-to-speech (TTS) model by distilling a large flow-matching teacher (IndicF5, 337M-parameter DiT) under a severe data budget (~17.6 hours). Training a small model from scratch on this much data fails outright. Instead we warm-start the student from the teacher by pruning depth only: keeping the teacher's width, text dimension, attention heads, and mel/text I/O fixed so all non-block tensors copy one-to-one, and retaining an evenly-spaced subset of transformer blocks. We first measure how much depth the teacher tolerates (it remains near-functional at -27% blocks but collapses past -50%), then descend gradually (22 -> 16 -> 12 -> 8 -> 6 blocks), re-fine-tuning after each prune, with each step gated by an objective ASR word-error-rate (WER) check. The resulting students reach WER 0.00 on unseen sentences at 249M and 190M parameters, and remain robust down to 131M; at 102M we observe a clear capacity cliff that we attribute to the data budget rather than the recipe. We also document two train/inference feature- and library-parity failures (mel filterbank and rotary-embedding library versions) that silently degrade audio, and a version-independent fix. The method yields a high-quality Hindi voice that runs in real time on a 6 GB laptop GPU. An independent 50-sentence FLEURS benchmark compares the released 190M student against its teacher and MMS-TTS-hin.