arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Halluscoring 2026:关于大语言模型幻觉检测与答案验证的首个共享任务

Halluscoring 2026: The first shared task on llms hallucination detection and answer verification

Aisha Alansari, Abdessalam Bouchekif, Ahmed Hasanaath, Salah Eddine Bekhouche, Malak Alkhorasani, Mohammed-En-Nadhir Zighem, Saad Ezzini, Hichem Telli, Hend Al-Khalifa, Muhammad Abdul-Mageed, Hadid Abdenour, Hamzah Luqman

arXiv 2609.38355首次发表:更新:

发表机构

King Fahd University of Petroleum and Minerals; Hamad Bin Khalifa University; University of the Basque Country; Imam Abdulrahman bin Faisal University; University of Biskra; University of British Columbia; Universiti Malaysia Kelantan; King Saud University(法赫德国王石油矿产大学; 哈马德·本·哈利法大学; 巴斯克大学; 伊玛目阿卜杜勒拉赫曼·本·费萨尔大学; 比斯克拉大学; 不列颠哥伦比亚大学; 马来西亚吉兰丹大学; 沙特国王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出HalluScoring 2026共享任务,评估阿拉伯语问答中幻觉检测与事实验证,包含四个子任务,13个团队参与,结果显示分布偏移下检测仍具挑战性。

AI 中文摘要

我们提出了HalluScoring 2026,这是一个在具有挑战性的泛化设置下,用于评估阿拉伯语问答中幻觉检测和事实验证的共享任务。该共享任务分为两个主要任务,每个任务包含两个子任务,共四个子任务。任务1评估二元幻觉检测,考虑对未见问题(子任务1.1)和未见大语言模型生成的回答(子任务1.2)的泛化能力。任务2将评估扩展到检测之外,要求系统从六个相关候选项中额外识别出正确的事实答案,涵盖伊斯兰知识(子任务2.1)和通用知识(子任务2.2)。该共享任务基于两个阿拉伯语数据集:HalluScore和HalluTruthQA。共有13个团队参与了该共享任务,其中10个团队提交了系统描述论文。任务1的结果表明,在分布偏移下幻觉检测仍然具有挑战性,获胜团队在子任务1.1和1.2中分别取得了0.772和0.767的AUC-ROC测试分数。对于任务2,在辅助评估下,获胜团队在伊斯兰知识和通用知识子任务中分别取得了0.882和0.857的分数。

英文摘要

We present HalluScoring 2026, a shared task for evaluating hallucination detection and factual verification in Arabic question answering under challenging generalization settings. Its four subtasks are organized into two tasks. Task~1 evaluates binary hallucination detection, considering generalization to unseen questions (Subtask 1.1) and responses generated by unseen LLMs (Subtask 1.2). Task 2 extends the evaluation beyond detection by requiring the systems to additionally identify the correct factual answer from six related candidates, covering Islamic knowledge (Subtask 2.1) and general knowledge (Subtask 2.2). The shared task is based on two Arabic datasets: HalluScore and HalluTruthQA. A total of 13 teams participated in the shared task, ten of which submitted system description papers. The results of Task 1 demonstrate that hallucination detection remains challenging. On the Task 1 test sets, the top-ranked systems achieved AUC-ROC scores of 0.7717 for Subtask 1.1 (REGLAT) and 0.7670 for Subtask 1.2 (NAMAA). Under assisted evaluation, the highest combined detection and answer-selection scores for Subtasks 2.1 and 2.2 were 0.8824 and 0.8565, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑