arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

搭建心智:优化多模态推理的潜在视觉目标表示

Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

Haoqiang Kang, Yinpeng Chen, Luyang Liu, Jesper Sparre Andersen, Abhijit Ogale, Baochen Sun, Lichan Hong, Ed H. Chi

arXiv 2608.19669首次发表:更新:

发表机构

Google DeepMind; UC San Diego(谷歌DeepMind; 加州大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对潜在推理框架的 SFT 阶段潜在表示对齐差、RL 阶段缺乏探索性潜在轨迹的局限,提出 Scaffolding Minds,学习专用搭建编码器与 RL 采样器的均值方差,在多基准任务上取得显著性能提升。

AI 中文摘要

潜在推理通过两阶段训练范式推动了多模态推理的发展:(1)在监督微调(SFT)阶段,将辅助图像编码为潜在 token,以教授视觉思维链;(2)在强化学习(RL)阶段,利用奖励反馈进一步优化这些潜在 token。本文中,我们识别出该框架在两个阶段各存在一个关键局限:其一,SFT 阶段通常依赖现成的视觉编码器对辅助图像进行编码,产生的潜在表示可能与下游推理任务对齐不佳;其二,现有 RL 方法仅通过确定性正则化处理潜在组件,虽能约束策略漂移,但无法生成用于探索的替代潜在轨迹。为解决这些局限,我们提出 Scaffolding Minds,该方法学习专用的搭建编码器以在潜在空间提供优化后的目标,并学习 RL 采样器的均值与方差。我们进一步表明,这两项改进具有互补性,共同相较于强基线取得了显著提升。实验表明,我们的方法在 FrozenLake 空间规划任务上比最强的潜在推理基线提升了 9.5%,在 32×32 网格地图上的提升幅度扩大至 19%,在 9 个视觉中心推理基准上平均提升了 5.2%。

英文摘要

Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage. In this paper, we identify two key limitations of this framework, one in each stage. First, the SFT stage typically relies on an off-the-shelf vision encoder to encode the helper image, yielding suboptimal latent representations that may not be well aligned with the downstream reasoning task. Second, existing RL methods treat the latent component only through deterministic regularization, which constrains policy drift but does not create alternative latent trajectories for exploration. To address these limitations, we propose Scaffolding Minds. Our approach learns a dedicated scaffolding encoder that provides an optimized target in latent space, and learns both the mean and variance of the RL sampler. We further show that these two improvements are complementary, together yielding substantial gains over strong baselines. Empirically, our method improves over the strongest latent reasoning baseline by +9.5 points on FrozenLake spatial planning, with the gain widening to +19 points on the 32x32 grids, and by +5.6 points on average across nine visual-centric reasoning benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑