Paradee:将Kokoro-82M蒸馏为8M参数的单语音文本转语音模型
Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model
浏览论文内容
中文总结 AI 辅助
该研究将Kokoro-82M蒸馏为8.07M参数的Paradee单语音TTS模型,参数量减少10倍,计算量降低15倍,通过两阶段训练和int8量化,实现单CPU线程25倍实时速度,UTMOS达4.41,并用锁相滤波器消除浊音噪声。
中文摘要 AI 辅助
我们将Kokoro-82M(一个广泛使用的、拥有54种语音的开源文本转语音模型)蒸馏为Paradee,一个拥有8.07M参数、仅能说其中一种语音的模型。Paradee保留了Kokoro的架构,但层宽大幅缩减,其两半部分分别针对冻结的教师模型进行训练。它的参数量减少了10倍,计算量需求降低了15倍。我们首先使用教师模型合成一个语料库,并保留其时长、音高、能量和音素特征。然后,我们训练一个小型文本侧网络来预测这些值,以及一个小型解码器,将教师模型保存的值转换为教师模型的音频,先使用频谱损失,然后使用对抗训练。最后,我们连接两半部分并将权重量化为int8。它无需对齐学习,也无需联合训练,并且可以在一台笔记本电脑上运行。以int8存储时,Paradee大小为8.5 MB,在单个CPU线程上运行速度比实时快25倍,UTMOS得分为4.41,而教师模型为4.52。学生模型最初保留了一点嗡嗡声,我们将其追溯到2至8 kHz之间的浊音相位。在合成后应用一个锁相滤波器可以去除大部分嗡嗡声,无需训练,也无需额外参数。代码、模型文件和音频样本可在https URL获取。
英文摘要
We distill Kokoro-82M, a widely used open text-to-speech model with 54 voices, into Paradee, an 8.07M-parameter model that speaks one of them. Paradee keeps Kokoro's architecture with much narrower layers, and each of its two halves is trained separately against the frozen teacher. It has 10x fewer parameters and needs 15x less compute. We first synthesize a corpus with the teacher and keep its durations, pitch, energy and phoneme features. We then train a small text side to predict these values, and a small decoder to turn the teacher's saved values into the teacher's audio, first with spectral losses and then adversarially. Finally, we connect the two halves and quantize the weights to int8. It needs no alignment learning and no joint training, and it runs on one laptop. Stored in int8, Paradee is 8.5 MB, runs 25x faster than real time on one CPU thread, and scores 4.41 on UTMOS against the teacher's 4.52. The student initially kept a slight buzz, which we trace to the phase of voiced speech between 2 and 8 kHz. A phase-locking filter applied after synthesis removes most of it, with no training and no extra parameters. Code, model files and audio samples are at https://github.com/sahilmahendrakar/paradee