发表机构
Singapore Management University(新加坡管理大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出令牌级感知基础优势估计(TPAE),利用视觉依赖性和预测熵的统计模式细化优势信号,提升多模态大语言模型推理的强化学习效果,在七个基准上优于强基线。
AI 中文摘要
基于可验证奖励的强化学习(RLVR)已提升了多模态大语言模型(MLLMs)的推理能力,然而现有框架依赖粗糙的序列级奖励信号,缺乏对多模态推理链中视觉基础步骤的细粒度监督。我们通过两个令牌级指标来研究这一差距:视觉依赖性(即令牌预测对输入图像特征的依赖程度)和预测熵。我们的实证分析揭示了两个关键发现:(1)与错误推理链相比,正确推理链在视觉基础增强时表现出明显更尖锐的熵降;(2)关键令牌(即其错误预测导致推理崩溃的令牌)在由正确链得出的视觉依赖性和预测熵的联合分布中是统计异常值。基于这些发现,我们提出了令牌级感知基础优势估计(TPAE),通过测量每个令牌与正确回滚的视觉熵模式的统计一致性来估计令牌级优势。TPAE利用这一粒度分数来调节序列级优势,产生可集成到各种RLVR框架中的细粒度监督信号。在七个基准上的广泛实验表明,TPAE始终优于领先的强基线,为多模态推理提供更稳定和高效的优化。代码可在以下https URL公开获取。
英文摘要
Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing frameworks rely on coarse, sequence-level reward signals that lack the fine-grained supervision over the visually-grounded steps within a multimodal reasoning chain. We investigate this gap through the lens of two token-level metrics: visual dependency (i.e. how much a token's prediction relies on the input image features) and predictive entropy. Our empirical analysis reveals two key findings: (1) correct reasoning chains exhibit a markedly sharper entropy reduction as visual grounding intensifies, compared to incorrect ones; (2) pivotal tokens, those whose misprediction triggers reasoning collapse, are statistical outliers in the joint distribution of visual dependency and predictive entropy derived from correct chains. Motivated by these findings, we propose token-level perception-grounded advantage estimation (TPAE), which estimates token-level advantages by measuring each token's statistical consistency with the vision-entropy patterns of correct rollouts. TPAE leverages this granular score to modulate the sequence-level advantage, producing a fine-grained supervision signal that can be integrated into various RLVR frameworks. Extensive experiments on seven benchmarks show that TPAE consistently outperforms leading strong baselines, yielding more stable and efficient optimization for multimodal reasoning. The code is publicly available at https://github.com/Zhihan72/TPAE.
CommentsAccepted by ACM MM 2026