arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SHROOM-Visions 2026概览:大型视觉语言模型幻觉检测共享任务

Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models

Raúl Vázquez, Aman Sinha, Chuyuan Li, Artem Shelmanov, Artem Vazhentsev, Claudio Savelli, Eduardo Calò, Emilio Raimond, Stella Frank, Hengyu Luo, Flavio Giobergia, Vincent Segonne, Lorenzo Vaiani, Jörg Tiedemann, Timothee Mickus

arXiv 2608.25662首次发表:更新:

发表机构

University of Helsinki; Politecnico di Torino; Université Bretagne Sud; University of Copenhagen; University Grenoble Alpes; University of Lorraine(赫尔辛基大学; 都灵理工大学; 南布列塔尼大学; 哥本哈根大学; 格勒诺布尔阿尔卑斯大学; 洛林大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文介绍了2026年在EMNLP 2026同期举办的SHROOM-Visions共享任务,该任务针对大型视觉语言模型的幻觉检测,依托SHEEP数据集开展,吸引27支团队提交600余个系统,最优系统较基线提升30-40个百分点。

AI 中文摘要

2026年,我们举办了SHROOM共享任务系列的第四届活动——SHROOM-Visions(视觉语言模型中幻觉及相关可观测过度生成错误的共享任务),该活动在与EMNLP 2026同期举办的UncertaiNLP研讨会上开展。继2024和2025年任务取得成功后,本次任务旨在通过针对大型视觉语言模型的模型无关检测任务解决幻觉问题。基于近期推出的、用于跨模型生成长期评估的SHEEP数据集,该任务邀请参与者检测并分类图像条件文本生成(视觉问答、图像字幕等)中的细粒度幻觉片段。评估采用涵盖中文、英文、法文、意大利文四种语言的五类幻觉分类体系。该共享任务在全球NLP社区引发了浓厚兴趣,共有27支团队提交了600余个系统结果。最优系统在四种语言上的字符级相关性平均得分为0.58、标签条件相关性得分为0.46、交并比(IoU)得分为0.51,较基线提升了30至40个百分点。

英文摘要

In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration \textbf{M}istakes in \textbf{Vision} language model\textbf{s}), which is hosted at the UncertaiNLP Workshop co-located with EMNLP 2026. Following the success of the 2024 and 2025 tasks, this time we aim to tackle hallucinations through a model-agnostic detection task focused on large vision-language models. Building on the recently introduced SHEEP dataset, designed for long-term evaluation across model generations, the task invites participants to detect and classify fine-grained hallucination spans in image-conditioned text generation (VQA, image captioning, etc.). The evaluation uses a five-class taxonomy of hallucinations spanning four languages: Chinese, English, French, and Italian. The shared task generated strong interest in the NLP community worldwide, with 27 teams contributing 600+ system submissions. The best systems achieve average scores of 0.58 in character-level correlation, 0.46 in label-conditioned correlation, and 0.51 in intersection-over-union (IoU) across four languages, outperforming the baselines by 30-40 points.

CommentsUnder review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑