arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

锚定关键:面向视觉基础多模态推理的双层学习框架

Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

Xinxin Song, Siyuan Li, Tingxiong Xiao, Jinli Suo

arXiv 2609.18057首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出PIVOT双层学习框架,通过自校准经验回放和视觉引导优势分配,解决RLVR在视觉基础多模态推理中的优化瓶颈,显著提升LVLMs推理能力。

AI 中文摘要

基于可验证奖励的强化学习(RLVR)显著提升了大视觉语言模型(LVLMs)的推理能力。然而,标准的在线策略RLVR算法在保留和强化视觉基础推理行为方面面临关键优化瓶颈:有价值的视觉基础推理轨迹在单次更新后被丢弃,而均匀的令牌优势分配阻止模型强化关键的感知或推理步骤。为弥合这一差距,我们提出PIVOT,一个双层学习框架,将策略优化锚定在信息丰富的视觉推理信号上。具体而言,PIVOT引入自校准经验回放机制,选择性地收集并回放视觉基础的历史经验作为策略优化的稳定参考锚点。在此基础上,我们进一步设计视觉引导的优势分配机制,根据令牌的局部视觉支持及其对下游推理的影响,为令牌分配额外的视觉感知优势。跨多个基准的大量实验表明,PIVOT在增强LVLMs的多模态推理能力方面取得了极具竞争力的性能。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capabilities of large vision-language models (LVLMs). However, standard on-policy RLVR algorithms face a critical optimization bottleneck in preserving and reinforcing visually grounded reasoning behaviors: valuable visually-grounded reasoning trajectories are discarded after a single update, while uniform token advantage allocation prevents the model from reinforcing critical perception or reasoning steps. To bridge this gap, we propose PIVOT, a dual-level learning framework that anchors policy optimization around informative visual reasoning signals. Specifically, PIVOT introduces a self-calibrated experience replay mechanism, which selectively collects and replays visually-grounded historical experiences as stable reference anchors for policy optimization. Building upon this, we further design a vision-guided advantage allocation mechanism to allocate additional vision-aware advantages to tokens based on their local visual support and impact on downstream reasoning. Extensive experiments across diverse benchmarks demonstrate that PIVOT achieves highly competitive performance in enhancing the multimodal reasoning capabilities of LVLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑