arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越词错误率:针对英语-约鲁巴语代码转换语音的ASR与音频语言模型的转换感知评估

Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech

Chibuzor Okocha, Christan Earl Grant

arXiv 2609.11786首次发表:更新:

发表机构

University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出转换感知评估框架,对11个ASR和音频语言模型在英语-约鲁巴语代码转换语音上进行测试,发现聚合词错误率掩盖转换行为,音频语言模型在转换局部指标上显著更优,并发布评估工具以支持可复现基准测试。

AI 中文摘要

自动语音识别(ASR)系统和音频语言模型(audio LMs)目前在单语基准上报告了较低的错误率,但它们在低资源、富含变音符号语言的代码转换语音上的行为仍缺乏充分表征。我们针对英语-约鲁巴语代码转换语音,对十一个现代系统(六个ASR模型和五个音频语言模型)进行了转换感知评估,使用了一个确定性的2000条话语评估集和一个共享的评分流水线。除了词错误率(WER)之外,我们报告了转换局部诊断指标:转换进入词元错误率(SETER)、窗口化转换点错误率、语言特定错误率以及变音符号不敏感的词错误率。我们的核心发现是,聚合词错误率掩盖了代码转换行为。在词错误率上表现最佳的系统(一个ASR模型)与一个领先的音频语言模型在统计上无显著差异,但该音频语言模型在每一个转换局部指标上均显著更优。在忠实系统之间,约鲁巴语词元识别崩溃(几乎所有系统的错误率均为0.97),而英语词元识别则好得多,且错误集中在转换进入约鲁巴语的节点。几个生成式音频语言模型作为精确转录器失败,产生翻译、冗长和提示泄漏,这些现象强烈依赖于提示。我们发布了清单、指标实现和评估脚本,以支持针对非洲代码转换语音的可复现、转换感知基准测试。

英文摘要

Automatic speech recognition (ASR) systems and audio language models (audio LMs) now report low error rates on monolingual benchmarks, but their behavior on code switched speech in low resource, diacritic rich languages remains poorly characterized. We present a switch aware evaluation of eleven modern systems (six ASR models and five audio LMs) on English Yoruba code-switched speech, using a deterministic 2000 utterance evaluation set and a shared scoring pipeline. Beyond word error rate (WER), we report switch localized diagnostics: a switch entry token error rate (SETER), windowed switch point error rates, language specific error rates, and a diacritic insensitive WER. Our central finding is that aggregate WER hides code switching behavior. The best system by WER (an ASR model) is statistically indistinguishable from a leading audio LM on WER, yet the audio LM is significantly better on every switch localized metric. Across faithful systems, Yoruba token recognition collapses (error 0.97 for almost all systems) while English tokens are recognized far better, and errors concentrate sharply at switches into Yoruba. Several generative audio LMs fail as exact transcribers, producing translation, verbosity, and prompt leakage that are strongly prompt dependent. We release manifests, metric implementations, and evaluation scripts to support reproducible, switch aware benchmarking for African code switched speech.

CommentsAccepted to IEEE Speech Language Tecnology

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑