arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DONDO:用于非洲语言的开放w2v-BERT语音识别基础模型

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

Paul Azunre, Naafi Ibrahim, Joel Budu, Lawrence Adu-Gyamfi

arXiv 2607.21540首次发表:更新:

发表机构

Khaya AI(Khaya人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出DONDO一族非洲语言语音识别基础模型,基于w2v-BERT 2.0构建,含单语和多语模型。介绍了两步微调及语言调节机制,多语模型平均字错误率10 - 13%,缩小与单语差距,模型开源可自由微调,覆盖众多非洲语言使用者。

AI 中文摘要

我们展示了DONDO,这是一族基于w2v-BERT 2.0自监督语音编码器构建的、开放且遵循宽松许可的非洲语言自动语音识别基础模型。DONDO包含21个单语模型和5个多语模型,涵盖加纳、塞拉利昂、尼日利亚、塞内加尔、肯尼亚和津巴布韦的27种语言变体。模型主要在宗教文本的朗读语音上进行微调,这些文本为缺乏转录音频的语言提供了广泛、许可清晰且拼写一致的覆盖。我们描述了一种两步(对于一个族为三步)学习率退火微调过程,首先以高学习率调整共享多语模型,然后退火以恢复并在某些情况下超过强大的单语基线。我们还描述了一种轻量级语言调节机制,在推理时将独热语言标识作为前缀帧序列注入声学特征,使单个多语检查点能转向目标语言。在五个多语族中,退火模型的平均字错误率达到10 - 13%,缩小了与单语模型的大部分差距,同时在单个检查点中覆盖多种语言。所有模型在Hugging Face KhayaAI组织下以Apache - 2.0许可(仅需署名)发布,以便他人可自由微调,包括商业用途。我们保守估计所涵盖语言的母语使用者约有一亿,若算上第二语言使用者则更多。

英文摘要

We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models and five multilingual models spanning twenty-seven language varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya and Zimbabwe. Models are fine-tuned primarily on read speech drawn from religious texts, which offer broad, license-clear and orthographically consistent coverage for languages that otherwise lack transcribed audio. We describe a two-step (and, for one family, three-step) learning-rate-annealed fine-tuning procedure that first adapts a shared multilingual model at a high learning rate and then anneals it to recover, and in several cases surpass, strong monolingual baselines. We further describe a lightweight language-conditioning mechanism that injects a one-hot language identity as a sequence of prefix frames prepended to the acoustic features, allowing a single multilingual checkpoint to be steered to a target language at inference. Across the five multilingual families the annealed models reach average word error rates (WER) of 10-13%, closing most of the gap to monolingual models while covering many languages in a single checkpoint. All models are released on the Hugging Face KhayaAI organisation under the Apache-2.0 license (attribution only) so that others may fine-tune them freely, including for commercial use. We provide a conservative estimate that the languages covered are spoken by on the order of one hundred million first-language speakers, and by substantially more when second-language use is included.

Commentsv2: added several co-authors; minor related-work revisions

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑