arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

追踪证据:通过视觉-语言推理实现忠实的令牌归因

Tracing the Evidence: Faithful Token Attribution Through Vision-Language Reasoning

Bowen Yuan, Danny Wang, Ruihong Qiu, Zijian Wang, Zi Huang

arXiv 2609.37656首次发表:更新:

发表机构

The University of Queensland(昆士兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有令牌归因方法在多模态场景下低估视觉证据及遗漏中间推理路径的问题,提出VTrace框架,通过中间推理追踪输入贡献并跨模态校准归因分数,实现忠实的多模态令牌归因。

AI 中文摘要

大型视觉-语言模型(LVLMs)展现出强大的推理能力,然而,支持生成响应的视觉和文本证据仍然难以识别。忠实的令牌归因通过分配分数来解释LVLM的响应,这些分数根据模型对图像和提示令牌的依赖程度对其进行排序,使得移除排名较高的令牌会导致生成响应的可能性更快下降。然而,现有的令牌归因方法主要针对基于文本的语言模型开发,我们的实证研究揭示了涉及复杂多模态来源时的两个挑战。首先,联合图像-文本归因相对于文本可能低估视觉证据,从而掩盖支持响应的图像区域。其次,视觉证据可能通过多个中间推理路径影响生成的响应,而现有方法仅追踪这些路径的有限子集,导致重要的视觉贡献被低估。受这些见解的启发,我们引入了VTrace,一个多模态令牌归因框架,它通过中间推理追踪输入贡献,并跨模态校准归因分数。VTrace构建成对归因以突出令牌特定的贡献,并以封闭形式聚合所有前向归因路径,以考虑直接和间接贡献。然后,跨模态校准使用从响应可能性变化估计的模态贡献来重新缩放图像和文本归因分数,从而实现输入令牌的统一排序。在六个视觉推理基准上对七个基线的评估证明了其优越的归因忠实度。项目页面:此https URL。

英文摘要

Large vision-language models (LVLMs) exhibit strong reasoning capabilities, yet the visual and textual evidence supporting the generated responses remains difficult to identify. Faithful token attribution explains an LVLM's response by assigning scores that rank image and prompt tokens by how much the model relies on them, such that removing higher-ranked tokens causes the likelihood of the generated response to drop more rapidly. However, existing token-attribution methods have been developed mainly for text-based language models, and our empirical study reveals two challenges when complex multimodal sources are involved. First, the joint image-text attribution can underrepresent visual evidence relative to text, obscuring the image regions supporting the response. Second, visual evidence may influence the generated response through multiple intermediate reasoning paths, while existing methods trace only a limited subset of these paths, causing important visual contributions to be underestimated. Motivated by these insights, we introduce VTrace, a multimodal token-attribution framework that traces input contributions through intermediate reasoning and calibrates attribution scores across modalities. VTrace constructs pairwise attributions that highlight token-specific contributions and aggregates all forward attribution paths in closed form to account for both direct and indirect contributions. Cross-modal calibration then rescales image and text attribution scores using modality contributions estimated from response-likelihood changes, enabling a unified ranking of input tokens. Evaluations against seven baselines across six visual reasoning benchmarks demonstrate the superior attribution faithfulness. Project page: https://vtrace-attribution.github.io/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑