基于可学习每通道能量归一化的时间包络前端用于Whisper儿童语音识别
A Temporal-Envelope Frontend with Learnable Per-Channel Energy Normalization for Whisper-Based Children's ASR
浏览论文内容
中文总结 AI 辅助
针对儿童语音识别,提出结合希尔伯特包络与可学习PCEN的时域前端,在MyST上使WER从13.16%降至11.08%,相对降低15.8%。
中文摘要 AI 辅助
时间包络携带对语音可懂度至关重要的线索,然而基于对数梅尔频谱图的ASR前端并未显式建模连续的子带包络结构。这一局限性对儿童语音尤为突出,因为儿童语音的高声学变异性要求鲁棒的特征表示。我们提出了一种模块化时域前端,利用梅尔间隔的窗函数sinc滤波器和希尔伯特变换将语音分解为子带包络,并采用可学习的每通道能量归一化(PCEN),与Whisper模型联合优化。在MyST儿童语音语料库上,系统性的消融实验确定了全频带窗函数sinc滤波器、希尔伯特包络、25 Hz平滑截止频率和可学习PCEN为最佳配置。在相同的Whisper-small微调设置下,该前端将词错误率(WER)从13.16%降至11.08%,相对降低了15.8%,优于对数梅尔基线,并在相同的清洗测试分割上超过了所评估的Kid-Whisper检查点。这些结果表明,时间包络表示和可学习的前端归一化是后端适配对儿童ASR的有效补充。
英文摘要
Temporal envelopes carry cues critical to speech intelligibility, yet ASR frontends based on log-mel spectrograms do not explicitly model continuous sub-band envelope structure. This limitation is particularly acute for children's speech, where high acoustic variability demands robust feature representations. We propose a modular time-domain frontend that decomposes speech into sub-band envelopes using mel-spaced windowed-sinc filters and the Hilbert transform, with learnable per-channel energy normalization (PCEN) jointly optimized with the Whisper model. On the MyST children's speech corpus, systematic ablations identify full-band windowed-sinc filters, Hilbert envelopes, a 25 Hz smoothing cutoff, and learnable PCEN as the best configuration. Under the same Whisper-small fine-tuning setup, the frontend reduces WER from 13.16% to 11.08%, a 15.8% relative reduction over the log-mel baseline, and outperforms the evaluated Kid-Whisper checkpoint on the same cleaned test split. These results show that temporal-envelope representations and learnable frontend normalization are effective complements to backend adaptation for children's ASR.