arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18856eess.AScs.AIcs.SDeess.SP

GrainSpeech:更少上下文,更多细节,实现紧凑语音合成

GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

Zitao Liang, Chang Gao

首次发表
浏览论文内容

中文总结 AI 辅助

GrainSpeech通过固定感受野卷积编码器和梅尔特定梯度方差监督,以264.8K参数实现紧凑语音合成,在MCU上达到17.9倍实时生成,质量媲美大模型。

中文摘要 AI 辅助

紧凑型声学模型面临质量与容量之间的严峻权衡。我们在此背景下研究两个因素:编码器上下文和梅尔频谱图监督。一项感受野缩放研究表明,将自注意力扩展到15个音素以上,在音高、能量或时长预测方面并未带来一致的改进。基于这一发现,我们引入了一个固定感受野的卷积编码器,分别将相应预测误差降低了36.0%、17.3%和3.4%。我们进一步表明,直接迁移图像域梯度方差监督可以恢复细粒度变化,但会降低预测质量,这促使我们提出一种针对梅尔频谱的特定公式,该公式包含轴特定梯度、重叠局部统计和对数域方差匹配。GrainSpeech仅包含264.8K个参数,在微控制器(MCU)上实现了17.9倍实时梅尔频谱生成,同时达到了与参数规模不到其1.5%的大得多模型相当的UTMOS分数。源代码和演示可在以下网址获取:此https URL。

英文摘要

Compact acoustic models face a challenging quality-capacity trade-off. We investigate two factors in this regime: encoder context and Mel-spectrogram supervision. A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction. Guided by this finding, we introduce a fixed-receptive-field convolutional encoder that reduces the respective prediction errors by 36.0%, 17.3%, and 3.4%. We further show that directly transferring image-domain gradient-variance supervision restores fine-scale variation but degrades predicted quality, motivating a Mel-specific formulation with axis-specific gradients, overlapping local statistics, and log-domain variance matching. GrainSpeech contains only 264.8K parameters and achieves 17.9x real-time Mel generation on a microcontroller (MCU), while attaining UTMOS scores comparable to substantially larger models with less than 1.5% of their parameters. Source code and demos are available at https://github.com/lab-emi/GrainSpeech.

发表机构

  • Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑