arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TTLab at AlexandriaX-2026:用于阿拉伯语机器翻译错误跨度检测与分类的微调表面标签器

TTLab at AlexandriaX-2026: A Fine-Tuned Surface Tagger for Arabic Machine-Translation Error-Span Detection and Classification

Ali Abusaleh, Bhuvanesh Verma, Alexander Mehler

arXiv 2609.29633首次发表:更新:

发表机构

Text Technology Lab (TTLab); Goethe University Frankfurt(文本技术实验室; 法兰克福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TTLab提交至AlexandriaX-2026,采用表面形式词级分类、焦点损失与方言阈值,MARBERTv2最佳,开发集40.8,测试集40.91,排名第3,稀有错误分类仍难。

AI 中文摘要

我们展示了TTLab在AlexandriaX-2026子任务3上的提交,该任务涉及阿拉伯语机器翻译错误跨度检测与分类。我们的系统将任务定义为基于表面形式的词级分类,保留字符偏移以确保与评估指标精确对齐。为处理严重的标签不平衡,我们采用带有类别加权的焦点损失以及方言特定的解码阈值。在六个阿拉伯语预训练编码器中,MARBERTv2在开发集和测试集上分别取得了40.8和40.91的最佳整体性能,在所有参赛队伍中排名第3。虽然我们的系统能有效定位错误跨度,但稀有错误类型的分类仍具挑战性,凸显了对尾部类别进行数据增强的必要性。代码可在https://this https URL上获取。

英文摘要

We present TTLab's submission to the AlexandriaX-2026 Subtask~3 on Arabic MT error span detection and classification. Our system frames the task as token-level classification over surface forms, preserving character offsets to ensure exact alignment with the evaluation metric. To handle severe label imbalance, we employ a focal loss with class weighting and dialect-specific decoding thresholds. Among six Arabic pre-trained encoders, MARBERTv2 achieves the best overall performance of 40.8 and 40.91 on the development and test set, respectively, ranking $\nth{3}$ out of all participating teams. While our system localizes error spans effectively, classification of rare error types remains challenging, highlighting the need for data augmentation for tail categories. The code is available at ${\href{https://github.com/ENTAILab/arabic-dialectal-mt-error-span-detection}{\faGithub~ TTLab at AlexandriaX-2026}$

CommentsAccepted at ArabicNLP 2026, shared task AlexandriaX-2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑