sanoTTS:通用微控制器上最小的实时神经TTS系统
sanoTTS: The Smallest Real-Time Neural TTS on a General-Purpose Microcontroller
浏览论文内容
中文总结 AI 辅助
本文提出sanoTTS,为通用微控制器上无需神经加速器的最小实时神经TTS,基于Piper/VITS教师模型蒸馏,规模小、速度快,在ESP32系列上实现实时语音生成,同时评估了其性能与局限。
中文摘要 AI 辅助
本文介绍了一种经审核的神经文本转语音(TTS)栈,可在通用微控制器上从音素ID生成22.05kHz的脉冲编码调制(PCM)信号。其部署的计算图拥有567008个参数,两个int8数据块占用679832字节。在ESP32-S3上,无需神经加速器,完整的时长-声学-逆短时傅里叶变换(inverse-STFT)路径生成4.54秒语音耗时1.02秒,实时率为0.22x;相同的便携式C核心在无浮点单元(FPU)的ESP32-C3上离线运行时实时率达5.72x。据所知,这是无需神经加速器、在通用微控制器上实时演示的最小完整音素到波形神经TTS计算图。我们从其Piper/VITS教师的条件变分自动编码器(conditional-VAE)目标中推导学生模型,并明确训练中使用的时长损失、潜在接口损失、波形损失、对抗损失及联合蒸馏损失。该模型的规模和速度存在可听代价:在未见过的文本上,从en_US-kristin-medium蒸馏出的嵌入式栈得分2.54 SCOREQ和2.80 UTMOS,而其教师的得分分别为4.68和4.42。另一套英文质量包使用更强的en_US-amy-medium教师,其拥有1454284个参数的帕累托点得分4.13 SCOREQ和4.10 UTMOS;拥有1834380个参数的变体得分4.16 SCOREQ。针对Kristin的受控容量研究表明,解码器而非输出表示是主要约束因素。两项评估失误影响了本研究:窄范围的模板化测试集使早期某学生模型的SCOREQ被高估了1.35,而聚合质量预测器未察觉在听觉测试和音素解析频谱探测中明显的擦音失败。校验和覆盖了所报告的模型数据块、运行时端口及基准向量。
英文摘要
This paper describes an audited neural text-to-speech stack that runs from phoneme IDs to 22.05-kHz PCM on general-purpose microcontrollers. Its deployed graph has 567,008 parameters, and its two int8 blobs occupy 679,832 bytes. On an ESP32-S3, the complete duration-acoustic-inverse-STFT path generates 4.54 s of speech in 1.02 s (0.22x real time) without a neural accelerator. The same portable C core runs offline at 5.72x real time on an FPU-less ESP32-C3. To our knowledge, this is the smallest complete phoneme-to-waveform neural TTS graph demonstrated in real time on a general-purpose microcontroller without a neural accelerator. We derive the students from the conditional-VAE objective of their Piper/VITS teachers and state the duration, latent-interface, waveform, adversarial, and joint-distillation losses used in training. The size and speed come with an audible cost: on unseen text, the embedded stack distilled from en_US-kristin-medium scores 2.54 SCOREQ and 2.80 UTMOS, compared with 4.68 and 4.42 for its teacher. A separate English quality package uses the stronger en_US-amy-medium teacher. Its 1,454,284-parameter Pareto point scores 4.13 SCOREQ and 4.10 UTMOS; a 1,834,380-parameter variant scores 4.16 SCOREQ. A controlled capacity study with Kristin identifies the decoder, rather than the output representation, as the main constraint. Two evaluation failures also affected the work: a narrow, templated test set overstated one early student's SCOREQ by 1.35, and aggregate quality predictors missed a sibilant failure that was evident in listening and in a phoneme-resolved spectral probe. Checksums cover the reported model blobs, runtime ports, and golden vectors.