arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18722cs.LGcs.CL

陈旧但稳定:用于稳定异步强化学习的陈旧自适应信赖域

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对异步强化学习中陈旧问题,提出陈旧自适应信赖域(SAT)方法,利用分离采样对数比率等技术识别高不匹配尾部并收缩PPO区间端点,实验表明该方法有效稳定异步强化学习,在特定设置下取得较好结果。

中文摘要 AI 辅助

异步强化学习通过将展开生成与优化解耦来提高吞吐量,但陈旧是由策略滞后、引擎延迟和专家混合路由等因素导致的不可避免的副产品。从信赖域的角度来看,这种不匹配至关重要。我们引入了陈旧自适应信赖域(SAT),它使用分离的采样对数比率作为实际的陈旧代理,通过基于陈旧的核缩放识别每批中的高不匹配尾部,并仅收缩名义近端策略优化(PPO)区间的符号选择端点。我们证明了相对于PPO的局部区间包含和逐点悲观性。我们在基于Qwen3 - 30B - A3B - Base构建的解耦异步强化学习设置中评估了SAT,结果表明SAT - GSPO w/ R3在滞后1时达到35.83,滞后8时达到34.79,SAT - GSPO在滞后1时达到34.17。自适应裁剪和路由重放分别针对不匹配尾部和路由不一致起到互补的稳定作用。总体而言,使裁剪区间与陈旧异质性对齐有效地稳定了异步强化学习。

英文摘要

Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but the resulting staleness is an inevitable byproduct, compounded jointly by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: in the finite-horizon improvement bound, training-inference divergence governs the approximation error, whereas PPO clipping only gates sampled outward updates and therefore acts as a sampled surrogate rather than a full-policy constraint. As a result, the high-staleness update can remain weakly controlled in exactly the asynchronous regime where stale rollouts matter most. We introduce the Staleness-Adaptive Trust Region (SAT), which uses the detached sampled log-ratio as a practical staleness proxy, identifies the high-mismatch tail within each batch through Staleness-based kernel function scaling, and contracts only the sign-selected endpoint of the nominal PPO interval using Effective contraction factors. This design preserves the baseline behavior on ordinary tokens, while making the update more conservative exactly on newly intercepted outward bands. We evaluate SAT in a fully decoupled asynchronous reinforcement learning setup built on Qwen3-30B-A3B-Base, leveraging SGLang as the inference engine and Megatron as the training pipeline. In this setting, SAT-GSPO w/ R3 attains the best observed AIME24 avg@8, reaching 35.83 at lag 1 and 34.79 at lag 8, while SAT-GSPO reaches 34.17 at lag 1. More broadly, the results indicate that aligning the clip interval with observed staleness heterogeneity is an effective way to stabilize the reported asynchronous regime.

补充信息

↑