arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PAIR:用于RLVR中自适应rollout分配的成对感知包含重加权方法

PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR

Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes

arXiv 2608.11368首次发表:更新:

发表机构

University of Wisconsin–Madison; University of Washington(威斯康星大学麦迪逊分校; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对RLVR中自适应rollout分配的统计不匹配问题,提出PAIR方法,通过成对对比图校正梯度估计,在Qwen3模型上提升准确率同时减少token使用量。

AI 中文摘要

带可验证奖励的强化学习(RLVR)将大部分计算资源用于生成长推理轨迹组。近期的分配器通过根据prompt、rollout或token的逐点难度或效用分配预算来降低该成本。我们发现存在统计不匹配问题:未截断的留一法组相对分数梯度并非独立点贡献之和,而是rollout对的二阶U统计量。因此,完成一个rollout会揭示其与所有其他已完成rollout的对比关系,而自适应端点选择会改变可观测的对项。我们提出PAIR(Pairwise-Aware Inclusion Reweighting,成对感知包含重加权),该方法将短rollout前缀视为顶点,将对梯度项视为对比图的边。仅前缀预测器可估计正确性与剩余token成本;凸设计在预期后缀token预算下选择正的继续概率;每个由已完成顶点诱导的边会按其记录的联合包含概率进行逆加权。在条件独立的在线策略rollout及未截断、未标准化目标下,所得估计器对完整候选对梯度是设计无偏的。在Qwen3-1.7B/4B上计算匹配的RLVR运行中,PAIR比最强的逐点分配器分别提升平均准确率1.2和1.4,同时比全组GRPO少生成51%和52%的token。冻结总体估计器审计证实,未加权自适应选择存在偏差,而成对包含校正可在匹配后缀成本下恢复完整对目标。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) spends most of its compute generating groups of long reasoning trajectories. Recent allocators reduce this cost by assigning budgets to prompts, rollouts, or tokens according to a pointwise notion of difficulty or utility. We identify a statistical mismatch: the unclipped leave-one-out group-relative score gradient is not a sum of independent point contributions, but a second-order U-statistic over pairs of rollouts. Completing one rollout therefore reveals contrast with every other completed rollout, and adaptive endpoint selection changes which pair terms are observable. We introduce PAIR (Pairwise-Aware Inclusion Reweighting), which treats short rollout prefixes as vertices and pair-gradient terms as edges of a contrast graph. A prefix-only predictor estimates correctness and remaining token cost; a convex design chooses positive continuation probabilities under an expected suffix-token budget; and each edge induced by completed vertices is inverse-weighted by its logged joint inclusion probability. Under conditionally independent on-policy rollouts and an unclipped, unstandardized objective, the resulting estimator is design-unbiased for the complete candidate-pair gradient. Across compute-matched RLVR runs on Qwen3-1.7B/4B, PAIR improves average accuracy by +1.2 and +1.4 over the strongest pointwise allocator while using 51% and 52% fewer generated tokens than full-group GRPO. A frozen-population estimator audit confirms that unweighted adaptive selection is biased, whereas pair-inclusion correction recovers the complete-pair target at matched suffix cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑