CARDAMOM:用于ASR的阿拉伯语微型方言语音数据集
CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR
AI总结:
本文提出Cardamom,一个覆盖21种阿拉伯语微型方言、约40小时的语音数据集,用于ASR细粒度评估与适配,基准测试显示适配后WER从43.47%降至35.21%,方言识别准确率达85.57%。
AI中文摘要:
我们提出了Cardamom,一个阿拉伯语微型方言语音数据集,旨在支持自动语音识别(ASR)系统的细粒度评估和适配。该数据集由熟悉所代表方言变体的母语者社区策展,包含约40小时的转录YouTube语音,覆盖埃及、约旦、黎巴嫩、毛里塔尼亚、巴勒斯坦和沙特阿拉伯的21种微型方言。每个片段都标注了一个或多个可操作的微型方言标签、代码切换信息以及话语级别的感知性别,从而能够分析传统国家级别标签所掩盖的次国家差异。我们描述了收集和标注过程,从语言学角度论证了微型方言清单的合理性,并在零样本和适配设置下对四个多语言ASR系统进行了基准测试。最强的零样本系统获得了43.47%的总体词错误率(WER),其中毛里塔尼亚和黎巴嫩方言的错误率特别高;在Cardamom上进行适配后,其WER降至35.21%。基于音频的识别实验进一步表明,这些标注提供了一个可学习的预测目标,一个专门的分类器在21类微型方言识别上达到了85.57%的准确率。Cardamom为研究局部方言变异和开发具有更广泛区域覆盖的阿拉伯语语音系统提供了资源。
英文摘要:
We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech spanning 21 micro-dialects across Egypt, Jordan, Lebanon, Mauritania, Palestine, and Saudi Arabia. Each segment is annotated with one or more operational micro-dialect labels, code-switching information, and utterance-level perceived gender, enabling analysis of sub-country variation that is obscured by conventional country-level labels. We describe the collection and annotation process, motivate the micro-dialect inventory linguistically, and benchmark four multilingual ASR systems in zero-shot and adapted settings. The strongest zero-shot system obtains 43.47% aggregate WER, with particularly high error rates on Mauritanian and Lebanese varieties; adaptation on Cardamom reduces its WER to 35.21%. Audio-based identification experiments further show that the annotations provide a learnable prediction target, with a dedicated classifier reaching 85.57% accuracy on 21-way micro-dialect identification. Cardamom provides a resource for studying localized dialectal variation and developing Arabic speech systems with broader regional coverage.