arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

策略滞后下的回滚复用:面向大语言模型强化学习的前缀归一化策略优化

Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning

Wenhao Zhang, Yibo Xie, Rui Wang, Jiahua Yang, Lei Jiang, Zibo Yang, Yawei Wang, Jiali Xu, jasperawang, Haoyang Long, Huan Xiong, alantzhao

arXiv 2608.01418首次发表:更新:

发表机构

Tencent; Harbin Institute of Technology; Jinan University; University of Science and Technology of China(腾讯; 哈尔滨工业大学; 暨南大学; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型强化学习中回滚复用导致的离策略问题,提出前缀归一化策略优化(PNPO),在4个更新周期时其数学推理性能优于GSPO,可降低更新批次需求。

AI 中文摘要

自回归回滚生成是大语言模型强化学习中的主要计算成本,将每个回滚批次复用至额外的学习者更新可摊销该成本,但随着学习者偏离行为策略,后续更新的离策略程度会不断提升。在某个 token 位置,精确的离策略校正必须同时考虑当前动作及其前缀的到达概率,累积重要性比率可实现该校正,但其乘积形式会产生难以处理的动态范围。我们研究前缀归一化策略优化(Prefix-Normalized Policy Optimization,PNPO),该方法用每个因果前缀上似然比的几何均值替代累积比率,在保留每个位置因果前缀依赖关系的同时压缩对数权重尺度。在受控的长上下文数学推理实验中,我们通过为每个回滚批次使用1或4个策略更新周期,诱导两种离策略 regime。PNPO在1个周期时未始终优于GSPO;在4个周期时,其在每个基准上达到观测到的最高Avg@32,三个独立选取的基准峰值的未加权均值为50.24,较GSPO高出3.00个百分点。在匹配的2400次更新预算下,4个周期的PNPO在150个回滚批次后达到最终宏Avg@32为49.66,与1个周期在600个批次后达到的49.56相当。这些结果为PNPO在训练进一步偏离策略时具备优势提供了初步证据。

英文摘要

Autoregressive rollout generation is a major computational cost in reinforcement learning for large language models. Reusing each rollout batch for additional learner updates amortizes this cost, but later updates become increasingly off-policy as the learner departs from the behavior policy. At a token position, exact off-policy correction must account for both the current action and the probability of reaching its prefix. The cumulative importance ratio provides this correction, but its product form can produce an unwieldy dynamic range. We study Prefix-Normalized Policy Optimization (PNPO), which replaces the cumulative ratio with the geometric mean of likelihood ratios along each causal prefix, preserving causal-prefix dependence at each position while compressing the log-weight scale. In controlled long-context mathematical reasoning experiments, we induce two off-policy regimes by using one or four policy-update epochs per rollout batch. PNPO does not consistently outperform GSPO with one epoch. With four epochs, it attains the highest observed Avg@32 on each benchmark; the unweighted mean of the three independently selected benchmark peaks is 50.24, 3.00 percentage points above GSPO. Under a matched 2,400-update budget, four-epoch PNPO reaches a final macro Avg@32 of 49.66 after 150 rollout batches, comparable to the 49.56 reached after 600 batches with one epoch. These results provide preliminary evidence that PNPO can be advantageous as training moves further off-policy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑