arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BoT-GRPO:通过词袋聚合实现推理的高效过程奖励强化学习

BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation

Yingxiang Yang, Weihang Xiao, Zhunxuan Wang, Joshua Flashner, Niresh Agarwal

arXiv 2610.09804首次发表:更新:

发表机构

Amazon AGI; Virginia Tech(亚马逊通用人工智能; 弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出BoT-GRPO,通过长度不变的词袋聚合实现令牌级奖励的高效过程监督,无需价值网络,在代码生成和数学推理任务上加速收敛并提升性能。

AI 中文摘要

强化学习现已成为大型语言模型中激发推理能力的核心手段,而在流行的算法组相对策略优化(GRPO)中,一次采样(rollout)中的每个令牌都获得相同的优势。我们探讨如何使过程监督高效化:在无需价值网络成本的情况下加速收敛并提升最终质量。我们提出词袋组相对策略优化(BoT-GRPO),通过一种长度不变的“词袋”聚合将GRPO扩展到令牌级奖励模型:它收集所有采样中的令牌级奖励,按源序列长度的倒数加权,并相对于加权组统计计算每个令牌的优势。BoT-GRPO无需评论家(critic-free),并且在令牌级奖励可用时,可作为GRPO的即插即用替代品。在React前端代码生成任务中,BoT-GRPO达到80%的编译率,比GRPO快最多1.9倍,并且比现代GRPO变体(GSPO、DAPO、PURE)收敛更快,同时达到更高的最终编译率和VLM评判的胜率。在第二个任务AIME数学推理中,BoT-GRPO在减半的步骤内,相对于GRPO实现了最多8.1%的绝对Pass@k增益。对于这两个任务,我们比较了算法在推理型与非推理型基础模型系列(Qwen2.5-3B、SmolLM3-3B、Phi-4-mini-reasoning)上的性能。我们的实验还为奖励模型本身提供了实用配方:奖励稳定性比丰富性更重要:干净、有界、稳定的细粒度信号持续加速学习,而更嘈杂的替代方案则停滞不前。

英文摘要

Reinforcement learning is now central to eliciting reasoning in large language models, while in the popular algorithm Group Relative Policy Optimization (GRPO) every token in a rollout receives the same advantage. We ask how to make process supervision efficient: accelerating convergence and improving final quality without the cost of value networks. We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics. BoT-GRPO is critic-free, and is a drop-in replacement wherever GRPO is used when token-level reward is available. On React front-end code generation, BoT-GRPO reaches $80\%$ compile rate up to $1.9\times$ faster than GRPO and converges faster than modern GRPO variants (GSPO, DAPO, PURE) while reaching higher final compile and VLM-judged win rates. On a second task, AIME mathematical reasoning, BoT-GRPO delivers absolute Pass@$k$ gains up to $8.1\%$ over GRPO in half the steps. For both tasks we compare the algorithm's performance on reasoning vs. non-reasoning base-model families (Qwen2.5-3B, SmolLM3-3B, Phi-4-mini-reasoning). Our experiments also yield a practical recipe for the reward model itself: reward stability matters more than richness: clean, bounded, stable fine-grained signals consistently accelerate learning where noisier alternatives stall.

CommentsPublished at the COLM 2026 Workshop on Efficient Reasoning

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑