arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29798cs.CLcs.AI

加纳语言青少年健康沟通中自动语音识别(ASR)的基准测试与领域自适应

Benchmarking and Domain Adaptation of Automatic Speech Recognition (ASR) for Adolescent Health Communication in Ghanaian Languages

Stephen E. Moore, Akwasi Asare, Mich-Seth Owusu, Paul Azunre, Joel Budu, Lawrence A. Adu-Gyamfi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究对三种加纳语言的ASR系统进行基准测试和领域自适应,通过微调Qwen3-ASR-0.6B显著降低WER,并部署KasaHealth应用验证,表明领域数据是关键制约因素。

中文摘要 AI 辅助

本文对三种加纳语言(Twi、Dagbani和Ewe)的青少年健康沟通中的自动语音识别(ASR)进行了端到端研究。工作分三个相互关联的阶段进行:首先,我们在通用领域的圣经语料库和青少年性与生殖健康(ASRH)领域ASR数据集上,使用字符错误率和词错误率(CER、WER)对五个ASR系统(三个特定语言的Wav2Vec2模型和两个多模态大语言模型Gemma 3n与Gemma 4)进行基准测试。其次,在基准测试的指导下,我们进行有监督的领域自适应:尽管Gemma 4是最强的零样本候选,但对其微调在计算上不可行,因此我们转向紧凑的Qwen3-ASR-0.6B,在大型加纳圣经语料库(约9万样本)上微调,并严格在留出的、人工收集的领域内音频上评估。微调降低了每种语言的WER,其中Ewe语最为显著(WER从109.3%降至64.8%,下降44.5个百分点;CER从65.1%降至24.9%)。第三,我们通过KasaHealth验证了这项工作,这是一个以语音为先的实时ASRH应用,部署在所有三种语言中,并辅以Senti-Check技术评估工具。KasaHealth接受了50名社区受访者的测试,实现了100%的聊天认可率、72%的良好或优秀翻译评分以及92%的推荐率,同时揭示了最制约实际应用的领域差距。在三个阶段中,证据汇聚于一点:对于这些语言,制约因素是经过验证的领域内数据,而非模型能力或计算资源。

英文摘要

This paper presents an end-to-end study of automatic speech recognition (ASR) for adolescent health communication in three Ghanaian languages (Twi, Dagbani, and Ewe). The work proceeds in three connected stages; First, we benchmark five ASR systems (three language-specific Wav2Vec2 models and two multimodal LLMs, Gemma 3n and Gemma 4) on a general-domain Bible corpus and a Youth Adolescent Sexual and Reproductive Health (ASRH) Domain ASR dataset, using Character and Word Error Rate (CER, WER). Second, guided by the benchmark, we perform supervised domain adaptation: although Gemma 4 was the strongest zero-shot candidate, fine-tuning it proved computationally infeasible, so we pivoted to the compact Qwen3-ASR-0.6B, fine-tuned on a large Ghana Bible corpus (~90k samples) and evaluated strictly on held-out human-collected in-domain audio. Fine-tuning reduced WER on every language, most dramatically for Ewe (WER from 109.3% to 64.8%, a drop of 44.5 pp; CER from 65.1% to 24.9%). Third, we validate the work through KasaHealth, a live voice-first ASRH application deployed in all three languages, complemented by Senti-Check, a technical evaluation harness. KasaHealth was tested by 50 community respondents and achieved a 100% chat-approval rate, a 72% Good-or-Excellent translation rating, and a 92% would-recommend rate, while surfacing the domain gaps that most constrain real-world use. Across all three stages the evidence converges: for these languages the binding constraint is validated in-domain data, not model capability or computation.

发表机构

  • University of Cape Coast(海岸角大学)
  • Ghana Natural Language Processing(加纳自然语言处理机构)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑