arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

COPC:异步大语言模型强化学习的耦合离策略校正

COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning

Zicheng Hu, Zhijian Zhou, Xuan Zhang, Yuchen Liu, Cheng Chen, Yuan Li, Qi Gu, Yan Feng, Hongyan Hao, Chao Qu

arXiv 2610.09597首次发表:更新:

发表机构

Fudan University; Meituan; East China Normal University(复旦大学; 美团; 华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对异步LLM强化学习中的策略与优势双重失配问题,提出耦合离策略校正(COPC)方法,通过协调策略侧和优势侧校正提升工具集成推理与搜索性能,并保持训练稳定性和低开销。

AI 中文摘要

异步强化学习通过将轨迹生成与优化解耦来加速大语言模型的后训练,但会在陈旧轨迹上进行训练。现有方法主要通过在actor目标中进行重要性比率控制来校正token级别的策略失配。我们证明仅靠这种“策略侧校正”是不够的:优势估计也会继承来自行为策略延续的失配,我们称之为“优势陈旧性”。我们推导了一般双通道actor更新的精确偏差和方差分解,揭示了策略权重误差与优势估计误差之间不可分离的耦合:它们的交互会产生乘性偏差项,而策略权重的平方会放大梯度方差中的优势不确定性。这促使我们提出假设:策略侧和优势侧的校正应当协调进行。我们引入了耦合离策略校正(COPC),这是一种actor-critic方法,结合了token级别的比率掩码和对TD残差进行双侧裁剪比率加权以进行回报和优势估计。跨陈旧性水平的联合参数扫描支持了这一假设:一个校正参数的效果取决于另一个参数,并可能随其反转。COPC在工具集成的数学推理和搜索中取得了最高报告性能,在每个设置中都优于最强报告的异步基线。它还提供了广泛的高性能参数区域和改善的训练稳定性。在搜索中,COPC在整个训练过程中保持稳定,而大多数评估的异步基线在训练后期崩溃。这些优势在64步策略陈旧性下依然存在。与异步PPO相比,COPC增加了极小的步时间开销,并保持了相对于同步PPO的1.7倍步时间加速。

英文摘要

Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective. We show that this \emph{policy-side correction} alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term \emph{advantage staleness}. We derive exact bias and variance decompositions for a general two-channel actor update, revealing nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in gradient variance. This motivates the hypothesis that policy- and advantage-side correction should be coordinated. We introduce Coupled Off-Policy Correction (COPC), an actor--critic method combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals for return and advantage estimation. Joint parameter sweeps across staleness levels support this hypothesis: the effect of one correction parameter depends on, and can reverse with, the other. COPC achieves the highest reported performance on tool-integrated mathematical reasoning and search, outperforming the strongest reported asynchronous baseline in each setting. It also offers a broad high-performing parameter region and improved training stability. In search, COPC remains stable throughout training, while most evaluated asynchronous baselines collapse late in training. These gains persist at 64-step policy staleness. COPC adds minimal step-time overhead over asynchronous PPO and retains a $1.7\times$ step-time speedup over synchronous PPO.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑