arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21811cs.AIcs.LG

提示、评论者与教师:视觉-语言数学推理中稀疏奖励强化学习的先验注入

Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning

Qiqian Fu

中文总结 AI 辅助

该研究针对视觉-语言数学推理的稀疏奖励强化学习问题,注入文本、分布、价值三类先验训练11种方法,发现先验有效传递可提升性能,还揭示了域内子集与跨域迁移的相关性反转现象,提出HL-Gauss损失替换MSE损失可提升14.4个百分点的准确率。

中文摘要 AI 辅助

视觉-语言数学推理的强化学习在稀疏奖励条件下性能受限:在包含20830道视觉数学题的数据集上,Qwen2-VL-2B的推理结果中仅3.6%的回合正确,85%-97%的GRPO(生成式强化学习优化算法)回合组完全错误,无法提供梯度。我们在该场景下于相同条件下训练了11种方法,每种方法注入不同类型的先验:文本先验(参考解决方案提示)、分布先验(基于7B教师模型的同策略蒸馏)、价值先验(经预训练的评论者,采用MSE或HL-Gauss分类损失)。先验仅在有效传递给策略时发挥作用:6种先验有效传递给策略的方法,在汇总的域内指标及跨域迁移(DynaMath数据集)上,与其余5种方法(无先验基线、4种先验被教师限制、门控或因评论者参数错误丢失的方法)完全分离,无重叠。核心发现与评估相关:域内数据集中长期用作项目通用分布检查的子集,与真实跨域迁移呈负相关(斯皮尔曼相关系数ρ=-0.74,n=11种方法,置换检验p=0.011),而最难的域内子集与跨域迁移密切正相关(ρ=+0.89,p<0.001)。我们将这种反转归因于接近随机的多项选择子集,该子集奖励模型不做改变;在该子集上,跨域表现最佳的方法表现平平,最差的方法看似是冠军。在这些方法中,提示引导的探索(而非UFT的辅助损失)驱动了提示带来的性能提升,将评论者的MSE损失替换为HL-Gauss交叉熵可使域内准确率提升14.4个百分点。所有准确率均为盲判结果,采用配对精确检验。

英文摘要

Reinforcement learning for vision-language math reasoning starves under sparse reward: on a pool of 20,830 visual-math problems where Qwen2-VL-2B answers 3.6% of rollouts correctly, 85-97% of GRPO rollout groups are entirely wrong and contribute zero gradient. We train eleven methods under identical conditions in this regime, each injecting a different prior: text (reference-solution hints), distribution (on-policy distillation from a 7B teacher), and value (a value-pretrained critic with an MSE or HL-Gauss categorical loss). A prior helps exactly when it is delivered: the six arms whose prior effectively reaches the policy separate with no overlap from the remaining five -- the no-prior baseline and four arms whose prior is teacher-capped, gated away, or lost to a mis-parameterized critic -- both on the pooled in-domain metric and on cross-domain transfer (DynaMath). The central finding, however, concerns evaluation: one slice of the in-domain pool -- long used as this project's general-distribution check -- anti-correlates with genuine cross-domain transfer (Spearman rho = -0.74, n = 11 arms, permutation p = 0.011), while the hardest in-domain slice predicts it closely (rho = +0.89, p < 0.001). We attribute the inversion to a near-chance multiple-choice subset that rewards models for not having changed; read through it, the best cross-domain method looked mediocre and the worst looked like the champion. Among the methods, hint-guided exploration -- not UFT's auxiliary loss -- drives hint gains, and replacing the critic's MSE loss with HL-Gauss cross-entropy is worth +14.4 points in-domain. All accuracies are blind-judged, with paired exact tests.

↑