发表机构
Qatar Computing Research Institute(卡塔尔计算研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了包含286名7-18岁说话者的151.72小时阿拉伯语L1/L2朗读语料库AraYoungVoices,基准测试显示L2语音更难识别,年龄特定微调改善匹配组,联合微调平衡各人群,且ASR假设更接近标准提示。
AI 中文摘要
最先进的自动语音识别(ASR)系统主要针对母语成人语音,导致在儿童、青少年和第二语言(L2)说话者上存在显著的性能差距。我们推出了AraYoungVoices,一个包含286名7至18岁说话者、总时长为151.72小时的阿拉伯语朗读语音语料库,由AraKids(7-12岁)和AraTeens(13-18岁)两个子集组成。该语料库包含146名阿拉伯语母语(L1)说话者和140名第二语言(L2)说话者,其中母语说话者涵盖埃及、海湾、黎凡特和北非方言背景,L2说话者则代表来自美洲、亚洲、非洲和欧洲的多样化语言背景。我们在零样本和微调设置下,使用未见说话者与未见提示(USUP)以及未见说话者与已见提示(USSP)评估,对四个预训练ASR模型进行了基准测试。结果表明,L2语音比L1语音更具挑战性,最大的错误主要出现在较年轻的L2说话者上。针对特定年龄的微调改善了匹配的年龄组,而联合微调则在各人群之间提供了更强的平衡。ASR假设也始终更接近标准朗读提示而非逐字转录,尤其是在L2语音中,这表明朗读偏差被部分规范化。
英文摘要
State-of-the-art ASR systems primarily target native adult speech, leading to substantial performance gaps for children, adolescents, and L2 speakers. We introduce AraYoungVoices, a 151.72-hour Arabic read-speech corpus from 286 speakers aged 7--18, comprising AraKids (7--12) and AraTeens (13--18). The corpus includes 146 native Arabic (L1) and 140 second-language (L2) speakers, with native speakers spanning Egyptian, Gulf, Levantine, and North African dialectal backgrounds and L2 speakers representing diverse linguistic backgrounds across the Americas, Asia, Africa, and Europe. We benchmark four pretrained ASR models under zero-shot and fine-tuned settings using unseen-speaker-$\&$-unseen-prompt (USUP) and unseen-speaker-$\&$-seen-prompt (USSP) evaluations. Results show that L2 speech remains substantially more challenging than L1 speech, with the largest errors observed mainly for younger L2 speakers. Age-specific fine-tuning improves the matched age group, while joint fine-tuning provides a stronger balance across populations. ASR hypotheses are also consistently closer to the standard reading prompt than to the verbatim transcription, particularly for L2 speech, suggesting partial normalization of reading deviations.