arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Almieyar:一个基于文化背景的多方言阿拉伯语语音识别基准

Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition

Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir, Malak Annan, Talha shahid javad allah rakha, Omar Al-Busaidi, Zineb El Kahla, Iheb Zouari, Essa Ahmed Abou Jabal, Ahmed Ezzat, Hind AL-Merekhi, Aisha Hamad M A Al-Naimi, Hadi Wazni, Bushra Alnajjar, Omar Amin, Haya Al-Thani, Houssam Eddine-Othman Lachemat, Marwa Elwakedy, Sundus Abdulmalik Al Nahari, Elahe Zahiri, Osamah Sarraj, Raghad Mousa, Mckeen Assi, Ahd Al Jumah, Heyam Salman, Alhanouf Abdulraqib, Sara Benoumhani, Alia Hamwi, Ayaat Al-Yasseri, Rim Ibrahim Ghazal, Lamia Ben hiba, Mohamed Eltabakh, Fatima Al-Raisi, Yassine El Kheir, Mohammed Abdulrahman, Hamdy Mubarak, Ayah Hashem, Lefkir Meriem, Ehsaneddin Asgari

arXiv 2609.35564首次发表:更新:

发表机构

QCRI, HBKU; Qatar University; UDST; Algo AI; UCL; University of Tripoli; AUC; KAUST; CMU-Q; KFUPM; Alfaisal University; Damascus University; Princeton University; ENSIAS, Mohammed V University; Sultan Qaboos University; DFKI; University of Waterloo; USTHB(卡塔尔计算研究所,哈马德·本·哈利法大学; 卡塔尔大学; 多哈科技大学; Algo AI; 伦敦大学学院; 的黎波里大学; 开罗美国大学; 阿卜杜拉国王科技大学; 卡内基梅隆大学卡塔尔分校; 法赫德国王石油与矿产大学; 费萨尔大学; 大马士革大学; 普林斯顿大学; 穆罕默德五世大学国家计算机科学与系统分析学院; 苏丹卡布斯大学; 德国人工智能研究中心; 滑铁卢大学; 阿尔及尔科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ALMIEYAR,一个覆盖17种阿拉伯语方言、基于文化场景录音的ASR基准,零样本评测12个系统,发现GPT-4o-transcribe最优(WER 35.0%),并揭示WER/CER差距,推动联合报告。

AI 中文摘要

阿拉伯语语音技术长期以来主要聚焦于现代标准阿拉伯语,而数亿人使用的日常方言却未得到充分服务。我们提出了ALMIEYAR,一个基于文化背景的自动语音识别(ASR)基准,覆盖六大语系的17种阿拉伯语方言,完全由现有模型未见过的全新录音构建而成。方言社区协调员从10个主题中挑选了具有文化相关性的图像,母语者通过五种结构化场景描述这些图像,每种方言约50分钟(总计13.7小时)。我们对12个最先进的ASR系统进行了零样本基准测试,包括GPT-4o-transcribe、Voxtral-Mini-4B、Fanar-STT-LF、Whisper、SeamlessM4T-v2以及基于wav2vec2的模型。GPT-4o-transcribe实现了最低的总体词错误率(WER)35.0%,其次是Voxtral-Mini-4B、Fanar-STT-LF和Whisper-Large-v3,分别为41.1%、45.9%和49.5%,表明阿拉伯语方言社区仍存在大量剩余错误。不同方言组的性能差异显著,没有模型在所有组中表现一致最优。仅凭WER也掩盖了方言ASR的行为:基于wav2vec2的模型显示出较大的WER/CER差距,字符级一致性远高于词级准确性,这促使采用联合WER/CER报告。ALMIEYAR为基于文化背景的阿拉伯语ASR评估提供了统一基准,包括首个已发布的Ahwazi阿拉伯语基准。

英文摘要

Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unseen by existing models. Dialect-community coordinators selected culturally relevant images across 10 topics, and native speakers described them through five structured scenarios, yielding approximately 50 minutes per dialect (13.7 hours total). We benchmark 12 state-of-the-art ASR systems zero-shot, including GPT-4o-transcribe, Voxtral-Mini-4B, Fanar-STT-LF, Whisper, SeamlessM4T-v2, and wav2vec2-based models. GPT-4o-transcribe achieves the lowest overall WER at 35.0%, followed by Voxtral-Mini-4B, Fanar-STT-LF, and Whisper-Large-v3 at 41.1%, 45.9%, and 49.5%, respectively, indicating substantial remaining errors across Arabic dialect communities. Performance varies considerably across dialect groups, with no model performing uniformly best across all groups. WER alone also obscures dialectal ASR behaviour: wav2vec2-based models show large WER/CER gaps, where character-level agreement remains much higher than word-level accuracy, motivating joint WER/CER reporting. ALMIEYAR provides a unified benchmark for culturally grounded Arabic ASR evaluation, including the first published benchmark for Ahwazi Arabic.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑