SVR-R1:在强化学习中通过自我验证引导多模态推理
SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
SVR-R1是一个多轮强化学习框架,将模型自身验证转化为学习信号。通过GRPO和异步多轮展开框架实现,无需外部监督。在视觉语言推理基准上评估,大幅提高准确率,弥合推理时自我优化与强化学习训练交叉点,还将开源以促研究。
中文摘要 AI 辅助
我们介绍了自我验证推理器(SVR-R1),这是一个多轮强化学习框架,它将模型自身的验证转化为多模态推理的学习信号。对于每个查询,模型使用相同权重提出答案并给出二元自我判断(是/否)。“否”会触发二次思考;“是”或达到轮次上限则确定输出以计算基于结果的奖励。SVR-R1通过GRPO和异步多轮展开框架实现,无需外部监督或辅助评论家。我们在视觉语言推理基准上评估SVR-R1,结果表明它比强大的标准GRPO基线大幅提高了准确率。训练动态显示对验证的依赖减少——验证轮次减少但测试准确率更高,这表明随着策略内化自我修正并通过我们的框架选择最自信答案,验证和生成之间的差距缩小。SVR-R1弥合了推理时自我优化与视觉语言模型强化学习训练之间较少探索的交叉点,为引导多模态推理提供了一个简单而有效的方法。我们将开源SVR-R1以促进视觉语言模型的未来研究。
英文摘要
We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning. For each query, the model proposes an answer using the same weights, and issues a binary self-verdict (Yes/No). A 'No' triggers a second-chance rethink; a 'Yes,' or a turn cap, finalizes the output for computing the outcome-based reward. SVR-R1 is implemented with GRPO and an asynchronous multi-turn rollout framework and needs no external supervision or auxiliary critics. We evaluate SVR-R1 on vision-language reasoning benchmarks and show that it improves accuracy by a large margin over strong standard GRPO baselines. Training dynamics show decreasing reliance on verification-fewer verification turns, yet higher test accuracy-indicating that the gap between verification and generation narrows as the policy internalizes self-correction and chooses the most confident answer via our framework. SVR-R1 bridges the less explored intersection of inference-time self-refinement and RL training for VLMs, offering a simple yet effective recipe for bootstrapping multimodal reasoning. We will open-source \textbf{SVR-R1} to facilitate future research in VLMs.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Meta
机构由 AI 辅助整理,请以论文原文为准。