AI 中文总结
针对阿萨姆语语音识别因数据不足面临挑战的问题,提出基于Whisper模型的可控微调系统,在特定语料库训练,并采用优化训练管道。该微调模型在多项指标上显著优于零样本基线,提升了语音识别性能。
AI 中文摘要
为形态丰富的低资源语言(如阿萨姆语)开发自动语音识别具有挑战性,因为标注语音数据不足,预训练的Whisper模型在阿萨姆语语音识别任务中表现不佳。本文提出了一个基于Whisper的可控微调阿萨姆语ASR系统,在Mozilla Common Voice 24.0 - 阿萨姆语语料库上训练。针对资源受限环境实现了硬件感知优化训练管道,采用混合精度训练和梯度累积。微调模型显著优于零样本基线,在词错误率、字符错误率等指标上有大幅提升,语义评估也有显著改进,预测幻觉率和实时因子也得到改善。
英文摘要
Developing Automatic Speech Recognition (ASR) for morphologically rich, low-resource languages such as Assamese is challenging due to insufficient annotated speech data. The pretrained Whisper model performs poorly on Assamese speech recognition tasks. This paper presents a controlled, fine-tuned Whisper-based Assamese ASR system trained on the Mozilla Common Voice 24.0-Assamese corpus. A hardware-aware optimized training pipeline is implemented for resource-constrained environments, employing mixed-precision training and gradient accumulation on Tesla 4 Graphics Processing Units (T4 GPUs). The proposed fine-tuned model significantly outperformed the Zero-shot baseline, yielding Word Error Rate (WER), Character Error Rate (CER), Match Error Rate (MER), and Word Infomation Loss (WIL) of 43.17\%, 13.18\%, 43\%, and 64.81\%, respectively, achieving significant relative improvements of 78.26\%, 93.10\%, 57.0\%, and 35.19\% over the baseline. Semantic evaluation of the fine-tuned model also demonstrates notable improvement over a zero baseline, attaining Bilingual Evaluation Understudy (BLEU) and Metric for Evaluation of Translation with Explicit ORdering (METEOR) scores of 30.81 and 0.5262, respectively. Additionally, the predicted hallucination rate and Real-Time Factor (RTF) are substantially improved by 96.70\% and 32.38\%, compared to the zero-shot baseline.
Comments31 pages, 11 figures