AI 中文总结
针对潜在视觉推理中潜在标记对视觉证据响应弱的问题,提出ReaLVR方法,通过引入视觉证据监督,在多个模型上提升性能,并首次将视觉推理扩展到235B规模。
AI 中文摘要
潜在视觉推理(LVR)使多模态大语言模型(MLLMs)能够在连续潜在标记中执行中间计算,而不是将每个推理步骤都用文字表达。然而,与文本思维链(CoT)不同,潜在推理不可直接观察,这使得监督潜在标记所学内容变得困难。在这项工作中,我们首先对潜在标记行为进行了深入分析,并识别出一个潜在证据-信用差距:潜在标记对改变正确答案的图像扰动仅作出微弱响应。我们假设这一问题源于GRPO训练期间缺乏显式监督。这些发现表明,最终答案奖励对保留哪些视觉证据或如何在潜在标记间分配信用提供了过少的指导。为弥合这一差距,我们提出ReaLVR,它将视觉证据监督引入模型自身的自由运行潜在轨迹。ReaLVR对比正确和模型生成的错误答案以确定何处需要更强监督,并对比相关和不匹配的视觉证据以指定要保留的内容。在三个模型家族中,ReaLVR持续优于评估的LVR基线,在Qwen2.5-VL-7B上实现了最高的五任务平均63.7%。关键的是,我们是首个在潜在空间中扩展视觉推理的团队,表明我们的框架在前沿模型规模(高达235B)上持续带来稳健改进。进一步分析显示,潜在标记位置对问题更敏感,与相关视觉区域的对齐更强,以及最受关注的潜在标记对固定上下文的依赖性更大。
英文摘要
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent tokens learn. In this work, we first conduct a thorough analysis of latent-token behavior and identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer. We hypothesize that this issue stems from the lack of explicit supervision during GRPO training. These findings suggest that a final-answer reward provides too little guidance on what visual evidence to preserve or how credit should be assigned across latent tokens. To bridge this gap, we propose ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. Across three model families, ReaLVR consistently outperforms evaluated LVR baselines, achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B. Crucially, we are the first to scale visual reasoning in latent space, showing that our framework continues to deliver robust improvements at frontier model scales up to 235B. Further analyses show more question-sensitive latent-token positions, stronger alignment with relevant visual regions, and greater fixed-context dependence on the most attended latent tokens.
Comments39 pages. Project page: https://xixiaouab.github.io/projects/ReaLVR/