发表机构
Department of Applied Mathematics and Computer Science; Technical University of Denmark(应用数学与计算机科学系; 丹麦技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究母语背景与英语自动语音识别(ASR)性能的关系,发现母语与英语的距离会系统性影响ASR错误率,且多数评估模型的深层声学层存在基于母语的潜在表示空间分离,该关联统计显著(p<0.001)。
AI 中文摘要
尽管自动语音识别(ASR)模型近年来取得了显著进步,但不同说话人群体间仍存在性能差异,其中一类差异针对母语(L1)与英语语系距离较远的说话者。本文研究母语背景与英语ASR性能的关系,通过实证分析发现,说话者的L1距离与ASR错误率的相关性对英语语音存在系统性影响,其强度随数据集和模型变化。在采用Tweedie混合效应模型考虑数据集层面变异的后续分析中,该关联具有统计显著性(评估模型中p均<0.001)。此外,对潜在空间的分析显示,在多数评估架构的更深声学层中存在基于L1的空间分离现象。
英文摘要
While automatic speech recognition (ASR) models have achieved remarkable improvements in recent years, performance disparities persist across different speaker populations. One such disparity is for speakers whose first languages (L1) are from families distant from English. This paper investigates the relationship between first language background and English ASR performance. Through empirical analysis, we observe that the correlation between speakers' L1 distance and ASR error rates yields a systematic effect on English Speech, with its strength varying across datasets and models. This association is statistically significant in a follow-up analysis accounting for dataset-level variation in Tweedie mixed-effects models ($p<0.001$ across evaluated models). In addition, analysis of the latent space reveals a L1-based spatial segregation across deeper acoustic layers in the majority of evaluated architectures
Commentsto appear in EMNLP finding 2026