MaPP:一种用于数据高效RLVR的统一边缘后验预测框架
MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR
- Beihang University(北京航空航天大学)
- Zhongguancun Academy(中关村学院)
- Communication University of China(中国传媒大学)
- Nanyang Technological University(南洋理工大学)
- Lobachebsky University(罗巴切夫斯基大学)
- Peking University(北京大学)
- Hangzhou Innovation Institute of Beihang University(北京航空航天大学杭州创新研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MaPP 提出统一边缘后验预测框架,通过 Beta-Binomial 边缘化去噪优势估计并改进提示选择,在数学、规划和视觉几何任务上超越 GRPO,实现数据高效 RLVR。
AI中文摘要:
基于可验证奖励的强化学习(RLVR)提升了大型语言模型的推理能力,但 rollout 和策略更新带来了高昂成本。在线提示选择通过使用每个提示的贝叶斯后验来预测难度并优先选择信息量大的提示,从而提高效率。然而,现有方法忽视了从采样响应中提取学习信号的可靠性。在 GRPO 中,一个响应的优势既取决于其自身的结果,也取决于通过组归一化随机采样的同伴结果。我们的理论和实验分析表明,组构成的不确定性引入了构成噪声,这是一种不消失的方差分量,对梯度估计误差施加了不可约的下界,并损害了后续的提示选择。我们提出了 MaPP(边缘后验预测),一个用于数据高效 RLVR 的统一框架,该框架对响应级优势估计进行去噪,并使用共享的 Beta 后验改进提示选择。对于每个响应,MaPP 通过闭式 Beta-Binomial 边缘化,用构成不变的内在优势替代标准的组相对优势。由此产生的后验预测估计器的误差在后验集中时被证明会减小。利用相同的后验,MaPP 推导出一个不确定性感知的提示选择分数,在不增加额外 rollout 成本的情况下提高数据效率。在五个模型骨干上的数学、规划和视觉几何实验表明,MaPP 始终优于 GRPO 和强选择基线,在相同 rollout 预算下,相对于最强基线实现了高达 +2.45 的平均准确率提升,并设立了新的最先进水平。
英文摘要:
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but incurs substantial costs from rollouts and policy updates. Online prompt selection improves efficiency by using per-prompt Bayesian posteriors to predict difficulty and prioritize informative prompts. However, existing methods overlook how reliably learning signals are extracted from sampled responses. In GRPO, a response's advantage depends on both its own outcome and the randomly sampled outcomes of its peers through group normalization. Our theoretical and experimental analyses show that uncertainty in group composition introduces composition noise, a non-vanishing variance component that imposes an irreducible lower bound on gradient estimation error and impairs downstream prompt selection. We propose MaPP (Marginalized Posterior-Predictive), a unified framework for data-efficient RLVR that denoises response-level advantage estimation and improves prompt selection using a shared Beta posterior. For each response, MaPP replaces the standard group-relative advantage with a composition-invariant intrinsic advantage through closed-form Beta-Binomial marginalization. The resulting posterior-predictive estimator has an error that provably diminishes as the posterior concentrates. Using the same posterior, MaPP derives an uncertainty-aware prompt selection score to improve data efficiency without additional rollout cost. Experiments on mathematics, planning, and visual geometry across five model backbones show that MaPP consistently outperforms GRPO and strong selection baselines, achieving up to +2.45 average accuracy improvement over the strongest baseline under the same rollout budget and setting a new state of the art.