低资源僧伽罗语的自动语音识别:方法、挑战与未来方向的关键性综述
Automatic Speech Recognition for Low-Resource Sinhala: A Critical Review of Methods, Challenges, and Future Directions
浏览论文内容
中文总结 AI 辅助
本文首次对低资源僧伽罗语ASR进行关键性综述,分析从HMM到自监督模型的发展,比较相关语言工作,识别六个研究空白并提出未来研究路线图。
中文摘要 AI 辅助
低资源语言的自动语音识别(ASR)仍然是一个重大挑战。僧伽罗语是斯里兰卡的主要语言,约有1600万使用者,它充分说明了这一困难:黏着语形态、54个音素的音位库、主语-宾语-动词(SOV)句法以及稀缺的标注语音数据,限制了传统和现代ASR系统的性能。本文首次对僧伽罗语ASR研究进行了关键性综述,追溯了其从隐马尔可夫模型(HMM)到深度神经网络,再到自监督预训练模型(如wav2vec 2.0、XLS-R、Whisper和大规模多语言语音(MMS))的发展历程。我们从架构、训练数据、词错误率(WER)以及对真实世界声学条件的鲁棒性等方面,将现有的僧伽罗语系统与泰米尔语、马拉雅拉姆语和印地语的相关低资源ASR工作进行了比较,并评估了自监督学习和迁移学习作为对稀缺标注数据的应对策略。我们表明,大多数报告的WER由于语料库、数据划分和评分方式的差异而无法直接比较,并且文献中唯一受控的比较将18.1%的相对WER降低仅归因于语料库修正。我们还讨论了利用音系、句法和语义知识的上下文感知ASR。我们识别出六个研究空白:(1)缺乏覆盖多种方言和声学条件的大规模标注语料库;(2)僧伽罗语形态句法的上下文建模薄弱;(3)真实世界条件下的高WER;(4)缺乏标准化基准;(5)缺乏参数高效微调研究;(6)缺乏标注的僧伽罗语-英语语码混合语音资源。我们概述了一个解决这些空白的研究议程,旨在为研究僧伽罗语和其他形态丰富语言的研究人员提供路线图。
英文摘要
Automatic speech recognition (ASR) for low-resource languages remains a major challenge. Sinhala, the primary language of Sri Lanka with about 16 million speakers, illustrates the difficulty: agglutinative morphology, a 54-phoneme inventory, subject-object-verb (SOV) syntax and scarce annotated speech data limit both conventional and modern ASR systems. This paper presents the first critical review of Sinhala ASR research, tracing its development from Hidden Markov Models (HMMs) through deep neural networks to self-supervised pre-trained models such as wav2vec 2.0, XLS-R, Whisper and Massively Multilingual Speech (MMS). We compare existing Sinhala systems with related low-resource ASR work on Tamil, Malayalam and Hindi in terms of architecture, training data, word error rate (WER) and robustness to real-world acoustic conditions, and we assess self-supervised and transfer learning as responses to scarce labeled data. We show that most reported WERs are not directly comparable because they differ in corpus, data split and scoring, and that the only controlled comparison in the literature attributes an 18.1% relative WER reduction to corpus correction alone. We also discuss context-aware ASR that draws on phonological, syntactic and semantic knowledge. We identify six research gaps: (1) the lack of large annotated corpora covering multiple dialects and acoustic conditions; (2) weak contextual modeling of Sinhala morphosyntax; (3) high WER in real-world conditions; (4) the absence of standardized benchmarks; (5) the lack of parameter-efficient fine-tuning studies; and (6) the absence of annotated code-switched Sinhala-English speech resources. We outline a research agenda to address these gaps, intended as a roadmap for researchers working on Sinhala and other morphologically rich languages.
发表机构
- University of Moratuwa(莫拉图瓦大学)
机构由 AI 辅助整理,请以论文原文为准。