arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AlexandriaX 2026:阿拉伯语方言机器翻译首次共享任务

AlexandriaX 2026: The First Shared Task on Dialectal Arabic Machine Translation

Abdellah El Mekki, AbdelRahim A. Elmadany, Samar M. Magdy, Saad Ezzini, Mo El-Haj, Mustafa Jarrar, Zaid Alyafeai, Bernard Ghanem, Muhammad Abdul-Mageed

arXiv 2609.22796首次发表:更新:

发表机构

The University of British Columbia; King Fahd University of Petroleum and Minerals; Lancaster University; VinUniversity; Hamad Bin Khalifa University; King Abdullah University of Science and Technology(不列颠哥伦比亚大学; 法赫德国王石油矿产大学; 兰卡斯特大学; VinUniversity; 哈马德·本·哈利法大学; 阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对阿拉伯语方言翻译挑战,提出AlexandriaX 2026共享任务,含三个子任务,最佳系统在受限和非受限轨道分别达30.42和33.49 spBLEU,资源公开。

AI 中文摘要

尽管阿拉伯语语言技术近期取得了进展,但阿拉伯语方言机器翻译(MT)仍然具有挑战性,特别是因为有效的翻译不仅需要对语义内容进行建模,还需要对方言变异、对话上下文、说话者和受话者特征以及社会语言学适切性进行建模。此外,传统MT指标对方言系统产生的语言错误提供的洞察有限。我们提出了AlexandriaX 2026阿拉伯语方言MT共享任务,通过三个互补的子任务解决这些挑战:(1)涵盖13种阿拉伯语变体的上下文感知英语到阿拉伯语方言对话翻译,(2)涵盖六种阿拉伯语方言的金融领域跨方言阿拉伯语翻译,以及(3)使用跨五种阿拉伯语变体的语言学动机错误类别进行跨度级MT错误检测和分类。该共享任务吸引了子任务1的38个注册、子任务2的33个注册和子任务3的35个注册。十二个独特团队提交了系统描述论文,我们全部接受发表。子任务1的最佳系统在受限轨道上达到了30.42 spBLEU,在非受限轨道上达到了33.49 spBLEU。在子任务2中,最佳系统获得了28.40 spBLEU。在子任务3中,最佳系统达到了49.82的总体得分,优于24.29的基线。总体而言,三个子任务中领先系统的结果凸显了显式方言建模、上下文感知生成、检索和重排序以及可解释MT错误分析专门方法的好处。AlexandriaX 2026共享任务的所有资源均公开可用,包括数据、基线和评估代码,可在我们的项目页面获取:此https URL。

英文摘要

Dialectal Arabic machine translation (MT) remains challenging despite recent progress in Arabic language technologies, particularly because effective translation requires modeling not only semantic content but also dialectal variation, conversational context, speaker and addressee characteristics, and sociolinguistic appropriateness. Moreover, conventional MT metrics provide limited insight into the linguistic errors produced by dialectal systems. We present the AlexandriaX 2026 Shared Task on Dialectal Arabic MT, which addresses these challenges through three complementary subtasks: (1) context-aware English-to-Dialectal Arabic dialogue translation across 13 Arabic varieties, (2) cross-dialect Arabic translation in the financial domain covering six Arabic dialects, and (3) span-level MT error detection and classification using linguistically motivated error categories across five Arabic varieties. The shared task attracted 38 registrations for Subtask 1, 33 for Subtask 2, and 35 for Subtask 3. Twelve unique teams submitted their system description papers, all of which we accepted for publication. The best system on Subtask 1 achieved 30.42 spBLEU in the constrained track and 33.49 spBLEU in the unconstrained track. On Subtask 2, the top system obtained 28.40 spBLEU. On Subtask 3, the best system achieved an overall score of 49.82, outperforming the 24.29 baseline. Taken together, the results of the leading systems across the three subtasks highlight the benefits of explicit dialect modeling, context-aware generation, retrieval and reranking, and specialized approaches to interpretable MT error analysis. All the resources of AlexandriaX 2026 shared task are publicly available, including data, baselines, and evaluation code on our project page: https://alexandriax.dlnlp.ai.

CommentsTo Appear in ArabicNLP 2026, resources available in the following project page: https://alexandriax.dlnlp.ai

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑