导航稀疏证据:通过显式上下文选择与整合的智能体视觉RAG
Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation
查看机构详情
- Soochow University(苏州大学)
- Baidu Inc.(百度公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对视觉RAG中证据稀疏与组织不足的问题,提出SCoRE智能体框架,通过显式选择与整合证据,结合轨迹蒸馏和强化学习,提升答案准确性与可追溯性。
中文摘要 AI 辅助
视觉检索增强生成(VRAG)使模型能够通过检索相关页面图像作为视觉证据并对其内容进行推理,来导航和回答关于视觉丰富文档的查询。然而,有效利用这种视觉证据通常受到两个主要挑战的阻碍。首先,与答案相关的证据是稀疏的,可能集中在某一页的一小片区域,或分散在多个页面中。其次,现有的智能体方法通常基于原始探索轨迹或压缩的文本记忆来生成答案,而非基于一组显式组织的支持图像,这使得答案容易受到探索噪声的影响,并模糊了基于证据的推理轨迹。我们认为瓶颈不仅在于证据发现,还在于在答案生成之前对证据的保存和组织。我们提出SCoRE(用于稳健证据的选择与整合),一个用于显式证据选择和整合的统一智能体循环。在探索过程中,SCoRE仅保留与查询相关的观察及其来源指针,并维护一个文本账本,在保持视觉上下文受限的同时保留早期证据。在终止时,它重新加载引用的原始图像并整合视觉证据以进行回答,将其排列成逻辑顺序。这将最终推理与探索性试错解耦,同时通过索引化的声明到图像链接确保严格的视觉基础。为了实现这一统一展开的端到端优化,我们的训练范式将过滤后的冷启动轨迹蒸馏与证据感知的强化学习相结合,其奖励促进证据覆盖、整合紧凑性和答案正确性。
英文摘要
Visual Retrieval-Augmented Generation (VRAG) empowers models to navigate and answer queries about visually rich documents by retrieving relevant page images as visual evidence and reasoning over their content. However, effectively utilizing this visual evidence is usually impeded by two main challenges. First, answer-relevant evidence is sparse and may be concentrated in a small region of one page or dispersed across multiple pages. Second, existing agentic methods often generate answers based on raw exploration trajectories or compressed textual memories rather than an explicitly organized set of supporting images, making answers susceptible to exploration noise and obscuring the evidence-backed reasoning trace. We argue that the bottleneck lies not only in evidence discovery but also in its preservation and organization before answer generation. We propose SCoRE (Selection and Consolidation for Robust Evidence), a unified agent loop for explicit evidence selection and consolidation. During exploration, SCoRE retains only query-relevant observations and their source pointers in a maintained textual ledger, preserving earlier evidence while keeping the visual context bounded. At termination, it reloads the referenced original images and consolidates the visual evidence for answering, arranging it into a logical sequence. This decouples final reasoning from exploratory trial-and-error while ensuring strict visual grounding via indexed claim-to-image linkages. To enable end-to-end optimization of this unified rollout, our training paradigm combines filtered cold-start trajectory distillation with evidence-aware reinforcement learning, whose reward promotes evidence coverage, consolidation compactness, and answer correctness.