发表机构
University of Helsinki; Politecnico di Torino; Universiteit Utrecht; Université Bretagne Sud; University of Copenhagen; University Grenoble Alpes; University of Lorraine(赫尔辛基大学; 都灵理工大学; 乌得勒支大学; 南布列塔尼大学; 哥本哈根大学; 格勒诺布尔阿尔卑斯大学; 洛林大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对模型迭代快导致幻觉评估普适性不足的问题,构建多语言人类编写幻觉样本数据集,证实其可替代模型生成样本用于细粒度视觉语言幻觉基准测试。
AI 中文摘要
在模型迭代迅速的时代,如何让幻觉评估更具普适性?本文探究人类编写的幻觉样本是否可替代模型生成的幻觉,从而使检测基准不依赖特定模型。为此,我们构建了包含1600个人类编写样本(覆盖中文、英文、法语、意大利语四种语言),以及来自五个视觉语言模型的18400个样本的数据集,所有样本均采用细粒度跨度级标注方案进行幻觉标注。研究发现,人类编写样本能带来更高的一致性,且可对数据集内容进行更精准的控制,同时其分布与视觉语言模型生成的样本相似,能合理反映检测性能,表明人类数据是基于模型的幻觉基准的可行替代品。
英文摘要
In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make benchmarking detection independent of particular models. To this end, we construct a dataset of 1,600 human-written samples, spanning four languages (Chinese, English, French, Italian), and 18,400 samples from five vision-and-language models, all annotated for hallucinations using a fine-grained span-level labeling scheme. We find that human-written samples result in higher agreement and allow greater control of dataset contents, while remaining distributionally similar to samples derived from vision-and-language samples and providing a reasonable portrayal of detection capabilities - suggesting that human data is a viable substitute for model-based hallucination benchmarks.