arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12327cs.CLcs.AIcs.LG

用于尼泊尔语自动语音识别的多语言预训练模型对比分析

Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition

Suman Paudel, Sarbin Sayami

首次发表
浏览论文内容

中文总结 AI 辅助

该研究微调6种多语言预训练模型开展尼泊尔语ASR对比实验,发现语系邻近性可替代模型规模,CTC解码器效率更高,MMS-1B鲁棒性更强,还提供了首个标准化参考基准。

中文摘要 AI 辅助

多语言预训练模型名义上支持尼泊尔语,但尚无受控基准在单一微调协议下对其进行比较。我们在OpenSLR SLR54尼泊尔语语料库(约165小时)上对6种预训练模型进行微调,这些模型涵盖CTC自监督、自回归编解码及混合Conformer-CTC架构,包括XLSR-53、IndicWav2Vec、MMS-1B、Whisper-Medium、Whisper-Large-v3-Turbo和Conformer-Hi,采用相同的预处理、数据划分、优化器及与模型家族匹配的学习率调度。我们在3个独立测试集(OpenSLR、FLEURS、Common Voice)上对词错误率(WER)、字符错误率(CER)和实时因子(RTF)进行评估。Whisper-Large-v3-Turbo(WER为14.76%)与IndicWav2Vec(WER为14.89%)并列最优,尽管参数量相差9倍、预训练数据量相差40倍,这提供了直接的经验证据:预训练中的语系邻近性可替代原始规模,适用于尼泊尔语的领域内任务。在相同准确率下,CTC解码器的运行速度比自回归Whisper快达29倍,这在任何延迟预算下都将实际部署偏好转向了CTC。大规模多语言预训练模型MMS-1B在FLEURS上的域外性能下降最小(+12.55个百分点),表明规模提升的是鲁棒性而非峰值域内准确率。由此产生的基准为尼泊尔语自动语音识别(ASR)提供了首个标准化、多模型、感知效率的参考数值。

英文摘要

Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium, Whisper-Large-v3-Turbo, and Conformer-Hi) spanning CTC self-supervised, autoregressive encoder-decoder, and hybrid Conformer-CTC architectures, on the OpenSLR SLR54 Nepali corpus (~165 hours) using identical preprocessing, splits, optimizer, and family-matched learning-rate schedules. We evaluate Word Error Rate (WER), Character Error Rate (CER), and Real-Time Factor (RTF) on three independent test sets (OpenSLR, FLEURS, Common Voice). Whisper-Large-v3-Turbo (14.76% WER) and IndicWav2Vec (14.89% WER) tie at the top despite a 9x parameter gap and 40x pretraining-data gap, providing direct empirical evidence that language-family proximity in pretraining can substitute for raw scale for in-domain Nepali. CTC decoders run up to 29x faster than autoregressive Whisper at the same accuracy, flipping the practical deployment preference toward CTC under any latency budget. Massively multilingual pretraining (MMS-1B) yields the smallest out-of-domain degradation on FLEURS (+12.55 pp), indicating that scale buys robustness rather than peak in-domain accuracy. The resulting benchmark provides the first standardized, multi-model, efficiency-aware reference numbers for Nepali ASR.

补充信息

↑