arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24651cs.CVcs.CLcs.IR

无坐标或区域标签的视觉文档理解中的证据归因

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

Zhuchenyang Liu, Yao Zhang, Yu Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

研究视觉文档理解中无坐标或区域标签时的证据归因问题,通过对比坐标与语言接口,发现语言接口能提升证据召回率、降低幻觉率。基于此用引述和检索管道作训练框架,引入GRPO方法,提高了模型严格归因准确率,找到改善归因的实用路径。

中文摘要 AI 辅助

可靠的视觉文档理解需要模型将每个答案归因于支持它的证据区域。近期基准和系统通过坐标接口表达这一步骤,然而视觉语言模型常在此接口下出现归因幻觉。本文研究该失败是否部分受模型通过坐标表达能力的限制。在经过验证的双语CiteVQA子集上,将坐标接口与语言接口(模型仅输出文本引述证据,多模态检索器返回引述位置)对比,六个开放视觉语言模型参与。结果显示与坐标接口相比,证据召回率提高,幻觉率降低,答案质量变化小。基于此,利用相同引述和检索管道作为训练框架,引入GRPO方法,在无区域标签情况下训练模型更好地引述证据,提高了严格归因准确率。这些发现表明了在无坐标接口和无昂贵区域级监督情况下改善归因的实用途径。

英文摘要

Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the evidence regions in the document. Under this interface, vision-language models often fail to identify the right regions even when the answer is correct, a failure known as Attribution Hallucination. We present a study that investigates whether this failure is partially limited by what the model can express through coordinates. On a verified bilingual CiteVQA subset, we compare the coordinate interface with a language interface in which the model outputs only text, quoting its evidence verbatim, and a multimodal retriever returns the location of each quote as a page region proposed by a layout parser (tables and figures are quoted through their captions or notes); the comparison is repeated over six open vision-language models. Compared with the coordinate interface, evidence recall rises from at most 8 points to between 26 and 47 and the hallucination rate roughly halves, with little change in answer quality. Building on this comparison, we use the same quote-and-retrieve pipeline as a training scaffold: because region-level evidence labels are expensive to collect for long documents, we introduce a GRPO recipe whose reward is a judge's reading of the gold answer and crops of the retrieved regions, training the model to quote better evidence without any region labels and raising an 8B backbone's strict attributed accuracy from 22.4 to 33.8. These findings indicate a practical path to improve attribution"without a coordinate interface and without costly region-level supervision.

发表机构

  • Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

↑