AI 中文总结
本文针对现有检索增强框架无法动态推理视觉证据插入时机与位置的问题,提出自主智能体框架VIG-RL,将相关工作流建模为主动决策过程,经强化学习优化后实现SOTA性能。
AI 中文摘要
在知识密集型场景中,提供可靠的文本-图像交互响应需要可验证的图像定位(Verified Image Grounding,VIG),即精准整合检索到的真实视觉证据。现有检索增强框架大多依赖解耦的静态流水线,无法动态推理何时需要外部知识、应在何处按上下文插入视觉资源。为填补这一空白,本文提出VIG-RL,这是一种自主智能体框架,将搜索-选择-插入工作流建模为主动决策过程。在动态ReAct风格的循环中运行,VIG-RL通过强化学习优化,由复合奖励系统指导,该系统全面评估智能体的逐步工具执行及最终多模态对齐。大量评估表明,VIG-RL达到了新的SOTA,显著优于现有静态基线。
英文摘要
In knowledge-intensive scenarios, providing reliable interleaved text-image responses requires Verified Image Grounding (VIG): the precise integration of retrieved authentic visual evidence. Existing retrieval-augmented frameworks predominantly rely on decoupled, static pipelines, inherently failing to dynamically reason about when external knowledge is required and where visual assets should be contextually inserted. To bridge this gap, we propose VIG-RL, an autonomous agentic framework that formulates the search-selection-insertion workflow as an active decision-making process. Operating within a dynamic ReAct-style loop, VIG-RL is optimized via reinforcement learning, guided by a composite reward system that holistically evaluates the agent's step-by-step tool execution and final multimodal alignment. Extensive evaluations demonstrate that VIG-RL establishes a new state-of-the-art, significantly outperforming existing static baselines.