微调Whisper以实现巴尼瓦语的自动语音识别:一项初步研究
Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study
浏览论文内容
中文总结 AI 辅助
本研究对Whisper Small进行微调,构建巴尼瓦语ASR初始基准,证实多语言基础模型可适配资源匮乏的土著语言,为后续相关研究提供支撑。
中文摘要 AI 辅助
近年来,自动语音识别(ASR)技术借助大型多语言基础模型取得了显著性能,但多数进展仍集中于高资源语言,而土著语言持续面临语音资源与语言技术匮乏的问题。本研究针对巴尼瓦语(一种分布于巴西、哥伦比亚和委内瑞拉的阿拉瓦克语系土著语言),开展了将Whisper适配至该语言自动语音识别的初步研究。实验使用的语料库来自某语言文档项目,包含1373份经人工转录的录音,时长约0.54小时,主要由孤立单词和简短引出语句构成。采用监督学习对Whisper Small模型进行微调,并以词错误率(WER)和字符错误率(CER)作为评估指标。最优模型达到37.5%的WER和7.45%的CER,表明多语言基础模型可成功适配至资源极度匮乏的土著语言。该研究为巴尼瓦语自动语音识别建立了初始基准,并为后续涉及更大数据集、语言特定适配策略及后处理技术的研究奠定了基础。
英文摘要
Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through the use of large multilingual foundation models. However, most advances remain concentrated on high-resource languages, while indigenous languages continue to suffer from a lack of speech resources and language technologies. This work presents a preliminary study on the adaptation of Whisper for Automatic Speech Recognition in Baniwa, an indigenous Arawakan language spoken in Brazil, Colombia, and Venezuela. The experiments were conducted using a corpus of 1,373 manually transcribed recordings obtained from a linguistic documentation project. The corpus contains approximately 0.54 hours of speech and consists primarily of isolated words and short elicited utterances. The Whisper Small model was fine-tuned using supervised learning and evaluated using Word Error Rate (WER) and Character Error Rate (CER). The best model achieved a WER of 37.5% and a CER of 7.45%, demonstrating that multilingual foundation models can be successfully adapted to extremely low-resource indigenous languages. The results establish an initial baseline for Baniwa Automatic Speech Recognition and provide a foundation for future research involving larger datasets, language-specific adaptation strategies, and post-processing techniques.
发表机构
- University of Brasilia(巴西利亚大学)
机构由 AI 辅助整理,请以论文原文为准。