arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DocAttriBench:文档视觉问答中的答案溯源基准

DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering

Luca De Grandis, Silvia Cappelletti, William Raccagni, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

arXiv 2609.20574首次发表:更新:

发表机构

University of Modena and Reggio Emilia; University of Pisa(摩德纳和雷焦艾米利亚大学; 比萨大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对文档VQA答案溯源缺乏标注且构建成本高的问题,提出DocAttriBench基准及MAPPET方法,基于掩码困惑度自动生成元素级溯源数据,并评测多模态大模型,发现强模型仍难定位支持元素。

AI 中文摘要

文档视觉问答中的答案溯源仍是一个开放挑战:大多数基准缺乏溯源标注或提供质量有限的标签,而构建带溯源的数据集仍需昂贵的人工努力。我们提出了DocAttriBench(DAB),一个用于文档VQA中细粒度、元素级来源归因的大规模基准,将答案溯源到特定的布局元素,如文本块、表格和图像。为构建DAB,我们提出了一种基于掩码的困惑度派生归因方法(MAPPET),该方法结合文档布局和语言建模来识别每个答案最具信息量的元素。MAPPET通过掩码候选元素后测量困惑度的增加,并将答案归因于对模型置信度贡献最大的元素。将MAPPET应用于多个现有文档VQA数据集,生成了DAB,包含23.7万份文档和29.6万个带有元素级溯源的问答对。我们在DAB上对具备溯源能力的多模态大语言模型进行了基准测试,评估答案准确率、溯源准确率和整体答案质量。结果表明,虽然较大的模型通常能达到更高的答案准确率,但即使最强的模型也常常无法定位支持元素。DAB为开发可溯源、可验证且可信的文档VQA模型提供了一个可扩展的基准。数据集和代码可在该https URL获取。

英文摘要

Answer grounding in document visual question answering remains an open challenge: most benchmarks lack grounding annotations or provide limited-quality labels, while constructing grounded datasets still requires costly manual effort. We introduce DocAttriBench (DAB), a large-scale benchmark for fine-grained, element-level source attribution in Document VQA, grounding answers to specific layout elements such as text blocks, tables, and images. To build DAB, we propose a Mask-based Perplexity-Derived Attribution method (MAPPET) that combines document layout and language modeling to identify the most informative element for each answer. MAPPET measures the increase in perplexity after masking candidate elements and attributes the answer to the element contributing most to model confidence. Applying MAPPET to multiple existing Document VQA datasets yields DAB, with 237k documents and 296k question-answer pairs with element-level grounding. We benchmark grounding-capable multimodal LLMs on DAB, evaluating answer accuracy, attribution accuracy, and overall answer quality. Results show that while larger models generally achieve higher answer accuracy, even the strongest models often fail to localize the supporting elements. DAB provides a scalable benchmark for developing grounded, verifiable, and trustworthy Document VQA models. Dataset and code are available at https://aimagelab.github.io/DocAttriBench/.

CommentsBMVC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑