arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37349cs.CVcs.AI

TAEC:面向多步视觉RAG的轨迹感知证据协调

TAEC: Trajectory-Aware Evidence Coordination for Multi-Step Visual RAG

Yalun Wu, Bingzhou Wang, Boyang Wang, Peiying Wang, Shaojie He, Yunhan Wang, Shaozu Yuan, Jiawei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多步视觉RAG中证据利用退化问题,提出无需训练的轨迹感知证据协调框架TAEC,通过跟踪未解决需求指导证据选择与记忆保留,在多个基准上取得最优性能。

中文摘要 AI 辅助

多步视觉检索增强生成(RAG)通过反复检索视觉证据、更新中间状态并决定是继续搜索还是回答,来回答复杂问题。然而,检索到相关证据并不能确保其在推理轨迹中得到有效利用。随着多步推理的推进,冗余来源占据了上下文容量,而这些容量本可用于缺失的证据;与已解决需求或无效搜索相关的观察结果滞留在上下文中;视觉来源被重新访问时缺乏细粒度读取所需的细节。我们将推理轨迹中可用证据的损失称为轨迹级证据利用退化。为解决这一问题,我们提出了轨迹感知证据协调(TAEC),一种无需训练的框架,围绕未解决的回答需求协调证据使用。TAEC在共享轨迹状态中跟踪这些需求,以指导哪些证据进入上下文、如何保留累积的记忆,以及以何种细节级别检查视觉证据。在ViDoSeek、SlideVQA和MMLongBench-Doc上的统一评估协议下,TAEC在领先的无需训练的视觉RAG基线中取得了最佳整体性能,并在多个专有视觉语言模型上获得了最高的平均准确率。这些结果表明,将证据与不断演变的推理需求对齐,可改善多步视觉RAG中证据的使用。

英文摘要

Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate state, and deciding whether to continue searching or answer. Yet retrieving relevant evidence does not ensure its effective use throughout the reasoning trajectory. As multi-step reasoning progresses, redundant sources occupy context capacity needed for missing evidence, observations tied to resolved requirements or unproductive searches linger in context, and visual sources are revisited with insufficient detail for fine-grained reading. We term this loss of usable evidence over a reasoning trajectory trajectory-level evidence utilization degradation. To address it, we propose Trajectory-Aware Evidence Coordination (TAEC), a training-free framework that coordinates evidence use around unresolved answer requirements. TAEC tracks these requirements in a shared trajectory state to guide which evidence enters the context, how accumulated memory is retained, and at what level of detail visual evidence is examined. Under a unified evaluation protocol on ViDoSeek, SlideVQA, and MMLongBench-Doc, TAEC achieves the best overall performance against leading training-free visual RAG baselines, with the highest average accuracy across multiple proprietary vision-language models. These results demonstrate that aligning evidence with evolving reasoning needs improves evidence use throughout multi-step visual RAG.

发表机构

  • National University of Singapore(新加坡国立大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

↑