AfriSwitch:用于野外非洲语码转换语音识别的基准测试集
AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition
浏览论文内容
中文总结 AI 辅助
该研究针对非洲语码转换语音识别缺失基准的问题,构建AfriSwitch基准测试集,测试发现现有多语言ASR系统性能远逊于单语言表现,且非洲定向训练是提升性能的关键。
中文摘要 AI 辅助
语码转换在双语非洲对话中十分普遍,但大多数自动语音识别(ASR)系统假设输入为单语言,并在精心整理的单语言基准上进行评估。我们推出AfriSwitch,这是一个时长61.36小时的野外语码转换语音人工转录基准,涵盖16种非洲语言及语言变体,附带语码转换层级的英语片段标签、每句话的语码混合指数(CMI)和转换点数量。语料库统计显示,非洲语言的语码转换行为在两个基本独立的维度上差异显著:说话人转换的频繁程度,以及混合的均衡程度。没有单一标量能准确衡量一种语言的语码转换程度。对5个开源及商用多语言ASR系统进行零样本基准测试,结果显示其词错误率(WER)远高于同语言已发表的单语言结果,最优系统平均WER为35.93%,且无任何系统在任一种语言上的WER低于24%。针对非洲的定向训练,而非模型规模或名义上的语言覆盖度,最能预测系统性能。
英文摘要
Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning 16 African languages and language varieties, released with switch-level English span tags, perutterance Code-Mixing Index (CMI), and switch-point counts. Corpus statistics show that mixing behaviour varies widely across African languages along two largely independent axes: how often speakers alternate, and how balanced the mixture is. No single scalar captures how code-switched a language is. Benchmarking five open and commercial multilingual ASR systems zero-shot yields word error rates far above published monolingual figures for the same languages, with the best system averaging 35.93% WER and no system falling below 24% on any language. Africa-targeted training, not model scale or nominal language coverage, best predicts performance.
发表机构
- Intron Health
机构由 AI 辅助整理,请以论文原文为准。