arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21595cs.CV

扰动思路而非像素:用于视觉语言模型强化学习的潜在空间展开多样化

Perturb the Thought, Not the Pixels: Latent-Space Rollout Diversification for Reinforcement Learning of Vision-Language Models

  • Amazon Web Services(亚马逊网络服务)

机构由 AI 辅助整理,请以论文原文为准。

Michael Jerge, Joseph Pelczar, Justin Downes

AI总结:

该研究提出 NC-GRPO 方法,通过在 VLM 潜在空间注入噪声实现展开多样化,提升了其数学推理与幻觉鲁棒性,且仅需少量代码修改即可整合进 RLVR 流程。

AI中文摘要:

带可验证奖励的强化学习(RLVR)可提升视觉语言模型(VLM)的推理能力,在每个优化组内多样化展开过程能放大其增益。现有方法通过解码温度或像素空间图像失真实现多样化,本文探究是否应将扰动置于模型的潜在空间中。我们引入噪声对比型 GRPO(NC-GRPO),在每个展开组的半数样本中,向提示编码阶段的最后隐藏层注入经尺度校准的高斯噪声,使这些展开过程从偏移后的起始状态分支。尽管存在偏移仍能得出答案的分支会比被其脱轨的分支获得更强的强化,将分支点处的敏感性转化为策略梯度信号;目标函数、奖励及推理协议均保持不变。在基于 Geometry3K 训练的 Qwen2.5-VL-7B 上,NC-GRPO 在五个保留基准测试中,相较 vanilla GRPO 显著提升了分布外数学推理能力(合并 McNemar 检验 p ≤ 0.001),同时还提升了域内准确率与幻觉鲁棒性——而图像空间噪声在感知密集型基准测试上虽呈现更大的分布外平均表现,却在幻觉鲁棒性这一维度出现退化。机制消融实验表明,独立随机多样性而非噪声预算或方向是核心有效成分,噪声尺度研究则揭示了推理专业化与通用能力间的调节关系。NC-GRPO 被设计为模态无关的方法,可作为约 50 行的修改整合进标准 RLVR 流程的推理引擎中。

英文摘要:

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning ability of vision-language models (VLMs), and diversifying the rollouts within each optimization group amplifies its gains. Existing approaches diversify through decoding temperature or pixel-space image distortion; we ask whether the perturbation belongs in the model's latent space instead. We introduce Noise-Contrastive GRPO (NC-GRPO), which injects scale-calibrated Gaussian noise into the last hidden layer of the prompt-encoding pass for half of each rollout group, branching those rollouts from a displaced departure state. Branches that reach the answer despite the displacement are reinforced over those derailed by it, converting sensitivity at the branch point into policy-gradient signal; the objective, reward, and inference protocol are untouched. On Qwen2.5-VL-7B trained on Geometry3K, NC-GRPO significantly improves out-of-domain mathematical reasoning over vanilla GRPO across five held-out benchmarks (pooled McNemar $p \le 0.001$) while also improving in-domain accuracy and hallucination robustness -- the latter an axis on which image-space noise regresses even while posting a larger OOD average on perception-heavy benchmarks. Mechanism ablations indicate that independent stochastic diversity, not noise budget or direction, is the active ingredient, and a noise-scale study exposes a dial between reasoning specialization and general capability. NC-GRPO is designed to be modality-agnostic and integrates into a standard RLVR pipeline as a ~50-line change to the inference engine.

↑