arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ImageEval 2026:基于文化语境的阿拉伯语多模态评估

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir, Md Arid Hasan, Basel Mousi, Nadir Durrani, Fahim Dalvi, Zien Sheikh Ali, Erchin Serpedin, Hasan Kurban, Mustafa Jarrar, Shammur Absar Chowdhury, Firoj Alam

arXiv 2608.30475首次发表:更新:

发表机构

Texas A&M University; Qatar Computing Research Institute; Birzeit University; Hamad Bin Khalifa University; University of Toronto(德克萨斯农工大学; 卡塔尔计算研究所; 比尔宰特大学; 哈马德·本·哈利法大学; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ImageEval 2026共享任务含AynVQA和CRAI-Bench两项阿拉伯语多模态评估任务,14支团队参与,相关资源已发布,凸显文化语境多模态评估的挑战。

AI 中文摘要

我们介绍了ImageEval 2026共享任务的概述,该任务聚焦基于文化语境的阿拉伯语多模态评估。它包含两项任务:(i)AynVQA,涵盖英语口语视觉问答以及英语和现代标准阿拉伯语(MSA)中的图像基础幻觉检测;(ii)CRAI-Bench,评估文生图生成的文化准确性。共有14支团队参与了测试阶段,其中12支团队提交了系统描述论文。参与的系统采用了多种方法,包括零样本提示、视觉语言模型的微调、语音识别流水线、集成学习以及分数校准。我们描述了任务设置、数据集、评估流程和参与系统,并总结了不同赛道的主要结果。该共享任务的所有数据集和评估脚本均已向研究社区发布。该共享任务凸显了基于文化语境的多模态评估面临的挑战,尤其是针对阿拉伯语语音和图像-文本推理的情况。

英文摘要

We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination detection in English and Modern Standard Arabic (MSA), and (ii) CRAI-Bench, evaluating the cultural accuracy of text-to-image generation. A total of 14 teams participated in the test phase, with 12 teams submitting system description papers. Participating systems used a range of approaches, including zero-shot prompting, fine-tuning of vision-language models, speech-recognition pipelines, ensembling, and score calibration. We describe the task setup, datasets, evaluation procedure, and participating systems, and summarize the main results across the different tracks. All datasets and evaluation scripts from the shared task are released to the research community. The shared task highlights the challenges of culturally grounded multimodal evaluation, particularly for Arabic speech and image-text reasoning.

CommentsArabic LLMs, Multilingual, Multimodal, Shared Task

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑