arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人类能否“梦到”电子羊?用于细粒度视觉语言幻觉基准测试的人类编写样本

Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking

Timothee Mickus, Claudio Savelli, Eduardo Calò, Emilio Raimond, Stella Frank, Hengyu Luo, Flavio Giobergia, Vincent Segonne, Chuyuan Li, Aman Sinha, Lorenzo Vaiani, Jörg Tiedemann, Raúl Vázquez

arXiv 2608.01021首次发表:更新:

发表机构

University of Helsinki; Politecnico di Torino; Universiteit Utrecht; Université Bretagne Sud; University of Copenhagen; University Grenoble Alpes; University of Lorraine(赫尔辛基大学; 都灵理工大学; 乌得勒支大学; 南布列塔尼大学; 哥本哈根大学; 格勒诺布尔阿尔卑斯大学; 洛林大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对模型迭代快导致幻觉评估普适性不足的问题,构建多语言人类编写幻觉样本数据集,证实其可替代模型生成样本用于细粒度视觉语言幻觉基准测试。

AI 中文摘要

在模型迭代迅速的时代,如何让幻觉评估更具普适性?本文探究人类编写的幻觉样本是否可替代模型生成的幻觉,从而使检测基准不依赖特定模型。为此,我们构建了包含1600个人类编写样本(覆盖中文、英文、法语、意大利语四种语言),以及来自五个视觉语言模型的18400个样本的数据集,所有样本均采用细粒度跨度级标注方案进行幻觉标注。研究发现,人类编写样本能带来更高的一致性,且可对数据集内容进行更精准的控制,同时其分布与视觉语言模型生成的样本相似,能合理反映检测性能,表明人类数据是基于模型的幻觉基准的可行替代品。

英文摘要

In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make benchmarking detection independent of particular models. To this end, we construct a dataset of 1,600 human-written samples, spanning four languages (Chinese, English, French, Italian), and 18,400 samples from five vision-and-language models, all annotated for hallucinations using a fine-grained span-level labeling scheme. We find that human-written samples result in higher agreement and allow greater control of dataset contents, while remaining distributionally similar to samples derived from vision-and-language samples and providing a reasonable portrayal of detection capabilities - suggesting that human data is a viable substitute for model-based hallucination benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑