发表机构
Graz University of Technology; Medical University of Vienna(格拉茨技术大学; 维也纳医科大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出多教师蒸馏框架训练轻量流式编码器,融合SSL离散目标和EL微调连续特征,将EL词错误率从39.3%降至21.2%,Mel-Conformer实现最佳效率与精度平衡。
AI 中文摘要
自监督学习(SSL)改善了语音表示,但在电喉(EL)语音等病理领域中性能会下降,且SSL模型的计算开销限制了其在实时、设备端部署中的适用性。我们提出了一种多教师知识蒸馏框架,用于训练一个轻量级、流式内容编码器,使其能够泛化到健康(HE)和EL语音。两位教师被逐步蒸馏:一个冻结的SSL模型,提供来自HE语音的离散音素聚类目标;以及一个在EL上微调的语音识别模型,提供连续的瓶颈特征目标。通过下游语音识别评估,我们的方法将EL词错误率降低至21.2%,而最强的零样本SSL基线为39.3%。在基于因果卷积、Transformer、Conformer和Mamba的学生架构中,Mel-Conformer在EL准确性和计算效率之间取得了最佳组合。最终编码器包含21.9百万参数,在单个CPU核心上以ONNX Runtime运行,实时因子为0.30。
英文摘要
Self-supervised learning (SSL) has improved speech representations, yet performance degrades in pathological domains such as electrolaryngeal (EL) speech, and the computational footprint of SSL models limits their applicability in real-time, on-device deployment. We propose a multi-teacher knowledge distillation framework to train a lightweight, streaming content encoder that generalizes across healthy (HE) and EL speech. Two teachers are distilled progressively: a frozen SSL model providing discrete phonetic cluster targets from HE speech, and an EL-fine-tuned speech recognition model supplying continuous bottleneck feature targets. Evaluated via downstream speech recognition, our approach reduces the EL word error rate to 21.2%, compared to 39.3% for the strongest zero-shot SSL baseline. Among causal convolutional, Transformer, Conformer, and Mamba-based student architectures, a Mel-Conformer achieves the best combination of EL accuracy and computational efficiency. The final encoder contains 21.9,M parameters and runs at a real-time factor of 0.30 under ONNX Runtime on a single CPU core.
Comments7 pages, 4 figures, 4 tables. Accepted at IEEE Spoken Language Technology Workshop (SLT) 2026