发表机构
Princeton University; Maseno University(普林斯顿大学; 马塞诺大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在大型多语言自动语音识别中,通过系统实验设计,探讨语言相关性能否可靠预测跨语言迁移增益,发现仅语言相关性难以有效促进跨语言迁移。
AI 中文摘要
将自动语音识别扩展到低资源非洲语言受数据收集限制。利用语言相关性通过顺序适应相关辅助语言和低资源目标语言来增强跨语言迁移,虽在小模型中有改进,但在大型模型中效果不明。通过系统实验设计扩展到大型多语言ASR,发现仅语言相关性难以有效促进跨语言迁移。
英文摘要
Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale. A promising direction is to leverage the linguistic relatedness between a low-resource target language and languages previously seen by a model to reduce the volume of target-language data needed for effective adaptation. Although this approach has proven reliable for text-based models, its effectiveness in the speech domain remains contested. We employ a systematic controlled experimental design spanning six factors, two Africa-centric corpora, and four large ASR models, sequentially adapting on a related auxiliary language followed by the target to isolate whether linguistic relatedness reliably predicts cross-lingual transfer gains across these conditions. In every setting, pre-adaptation on related auxiliary languages yields no practically meaningful improvements once as little as one hour of target-language data is available, suggesting that relatedness alone may not reliably predict transfer gains in large multilingual ASR, or constitute an effective strategy for extending such models to low-resource languages.