arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过强化学习激发多模态推理智能体的自我验证能力

Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning

Vishwas Sathish, Viresh Ranjan, Xinliang Zhu, Arnab Dhua, Douglas Gray

arXiv 2609.08025首次发表:更新:

发表机构

University of Washington; Amazon(华盛顿大学; 亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SVRL框架,通过强化学习训练多模态智能体在推理中自我验证和过滤检索证据,并引入搜索惩罚与查询多样性奖励,仅用5000个样本微调Qwen-2.5-VL-7B即显著提升多跳VQA泛化与工具效率。

AI 中文摘要

推理智能体日益依赖网络搜索等外部工具来回答复杂查询。GRPO等强化学习(RL)微调算法已提升了纯文本语言模型的长链推理能力,尤其在编码和数学领域。然而,多模态智能体中可靠的工具使用仍具挑战性,因为模型必须解读文本和图像,同时整合带有噪声的检索证据,且通常仅在稀疏的结果级监督下进行,缺乏明确的验证信号。我们提出了基于强化学习的自我验证(SVRL)框架,这是一种仅使用RL的微调框架,训练多模态智能体在其自身的推理轨迹中验证并过滤检索到的证据,从而减少推理时对外部验证器的依赖。SVRL还引入了一种搜索感知惩罚项,用于抑制不必要的工具调用,以及一种查询多样性奖励,用于鼓励生成多样且结构良好的搜索查询,从而对何时搜索以及搜索什么提供细粒度的反馈。仅使用5,000个视觉问答示例对Qwen-2.5-VL-7B进行SVRL微调,便在多跳VQA泛化能力和工具效率方面取得了跨基准的一致提升。总体而言,SVRL缩小了紧凑型智能体与更大规模专有模型之间的差距,同时大幅降低了训练和推理成本。

英文摘要

Reasoning agents increasingly rely on external tools such as web search to answer complex queries. Reinforcement learning (RL) finetuning algorithms such as GRPO have improved long-form reasoning in text-only language models, particularly for coding and mathematics. Reliable tool use in multimodal agents, however, remains challenging because models must interpret text and images while integrating noisy retrieved evidence, often under sparse outcome-level supervision without explicit verification signals. We present Self-Verification via Reinforcement Learning (SVRL), an RL-only finetuning framework that trains multimodal agents to verify and filter retrieved evidence within their own reasoning traces, reducing reliance on external verifiers at inference time. SVRL also introduces a search-aware penalty that discourages unnecessary tool calls and a query-diversity reward that encourages diverse, well-formed search queries, providing fine-grained feedback on when and what to search. Finetuning Qwen-2.5-VL-7B with SVRL on only 5{,}000 visual question answering examples yields consistent gains in multi-hop VQA generalization and tool efficiency across benchmarks. Overall, SVRL narrows the gap between compact agents and much larger proprietary models while requiring substantially lower training and inference cost.

CommentsTo appear in ECCV 2026. 11 main pages. 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑