arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19361cs.CLeess.AS

米佐语自动语音识别(ASR)语音语料库:结合形态感知评估的Whisper与SraVaani 1.0微调

A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

Priyankoo Sarmah, Sanasam Ranbir Singh, Lalhmingmawia

AI总结:

本研究构建米佐语语音语料库,通过微调Whisper与SraVaani 1.0模型开发米佐语ASR系统,发现Whisper-large-v3结合形态感知评估WER仅7.22%,SraVaani 1.0经米佐语微调后性能显著提升。

AI中文摘要:

本研究针对低资源语言米佐语开发自动语音识别(ASR)系统,流程包括收集17.62小时语音数据、整理数据,并使用三款Whisper多语言模型及SraVaani 1.0印度多语言模型对米佐语ASR系统进行微调。Whisper-large-v3的常规词错误率(WER)最低,为18.08%;而形态感知评估得到的WER为7.22%。对SraVaani 1.0印度多语言模型进行零样本评估时,其WER为58.27%;针对米佐语的微调将其常规WER降至29.45%,形态感知WER降至17.93%。结果表明,Whisper模型即使适配未接触过的语言也能实现极低WER;相比之下,SraVaani 1.0在其多语言模型中支持米佐语,但使用精心整理的米佐语语音数据进行微调可大幅提升其性能。

英文摘要:

This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system with three Whisper multilingual models and with the SraVaani 1.0 Indic multilingual model. Whisper-large-v3 achieved the lowest conventional WER (18.08%), while morphology-aware evaluation yielded a WER of 7.22%. Zero-shot evaluation of the SraVaani 1.0 Indic multilingual model yielded a WER of 58.27%, while Mizo-specific fine-tuning reduced the conventional WER to 29.45% and the morphology-aware WER to 17.93%. The results demonstrate that the Whisper model can achieve a substantially low WER, even when adapted to an unseen language. In contrast, SraVaani 1.0 supports the Mizo language in its multilingual model; however, fine-tuning with carefully curated Mizo speech data substantially improves its performance.

↑