arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAFE-G:面向基于知识的视觉问答的结构感知忠实证据引导生成

SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering

Long Shu, Shuochen Liu, Wei Chen, Junda Lin, Zhi Zheng, Huijun Hou, Tong Xu

arXiv 2608.21796首次发表:更新:

发表机构

State Key Laboratory of Cognitive Intelligence; University of Science and Technology of China; NIO(认知智能国家重点实验室; 中国科学技术大学; 蔚来)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对KB-VQA现有方法难以捕捉结构关联、推理不忠实的问题,提出SAFE-G框架,通过混合搜索、图检索与证据对齐RL策略,在两个基准上准确率优于现有方法8.9%、3.5%。

AI 中文摘要

基于知识的视觉问答(KB-VQA)旨在回答需要对视觉内容之外的外部知识源进行推理的查询。当前方法通常融合多模态特征以检索外部信息,随后利用多模态大语言模型(MLLMs)从检索到的证据中推导答案。然而,这些方法往往难以捕捉复杂语境中的结构关联以有效过滤噪声,且经常无法确保推理过程严格忠实于检索到的证据。为应对这些挑战,我们提出SAFE-G,即结构感知忠实证据引导生成框架,该框架可实现精确的证据定位与可信的推理。具体而言,我们首先采用融合视觉与文本模态的粗粒度混合搜索以召回候选文档,随后实施结构感知的细粒度图检索,该检索可捕捉结构依赖关系以过滤噪声并定位精确证据。此外,我们引入带有基于证据的奖励的强化学习(RL)策略,仅当所选证据正确时才为正确答案分配信用。这种严格的对齐约束迫使模型将其响应锚定在检索到的上下文中,有效增强其通过多模态特征定位证据并进行忠实推理的能力。在Encyclopedic-VQA和InfoSeek基准上的大量实验表明,SAFE-G以8.9%和3.5%的优势优于现有方法,大幅提升了整体推理准确率。我们的源代码可在以下网址公开获取:this https URL。

英文摘要

Knowledge-based Visual Question Answering (KB-VQA) aims to answer queries that necessitate reasoning over external knowledge sources beyond the visual content. Typically, current methods fuse multimodal features to retrieve external information, subsequently leveraging Multimodal Large Language Models (MLLMs) to derive answers from the retrieved evidence. However, these methods often struggle to capture structural associations within complex contexts to effectively filter noise. Furthermore, they frequently fail to ensure that the reasoning process remains strictly faithful to the retrieved evidence. To address these challenges, we propose SAFE-G, a Structure-Aware Faithful Evidence-guided Generation framework, which enables precise evidence localization and trustworthy reasoning. Specifically, we first employ a coarse-grained hybrid search fusing visual and textual modalities to recall candidate documents, and subsequently implement a structure-aware fine-grained graph retrieval that captures structural dependencies to filter noise and pinpoint precise evidence. Moreover, we introduce a reinforcement learning (RL) strategy with an evidence-grounded reward that assigns credit to correct answers only when the selected evidence is correct. This strict alignment constraint compels the model to anchor its response in the retrieved context, effectively enhancing its capability to locate evidence via multimodal features and perform faithful reasoning. Extensive experiments on the Encyclopedic-VQA and InfoSeek benchmarks demonstrate that SAFE-G outperforms prior methods by a margin of 8.9% and 3.5%, substantially enhancing the overall reasoning accuracy. Our source code is publicly available at: https://github.com/MINE-USTC/SAFE-G.

Comments12 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑