arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解构离策略比率:用于异步强化学习的熵缩放信任区域

Deconstructing Off-Policy Ratios: Entropy-Normalized Trust Regions for Asynchronous Reinforcement Learning

Guanqun Zhao, Zijun Xie, Binbin Zheng, Yehan Yang, Jiafeng Lu, Aoqi Hu, Enlei Gong, Zeyu Chen

arXiv 2607.22186首次发表:更新:

发表机构

Baidu(百度)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究异步强化学习中离策略数据问题,提出熵缩放信任区域(ESTR)方法,通过令牌局部熵缩放离策略偏差,无需辅助前向传递等,在多任务和基准测试中优于现有异步方法,提升训练-推理一致性与速度。

AI 中文摘要

异步强化学习通过将生成轨迹与策略优化重叠来加速大语言模型的训练后优化,但由此产生的陈旧离策略数据会破坏优化并导致策略崩溃。现有方法通常仅根据重要性比率的大小保留或丢弃令牌,在令牌位置上统一应用相同的阈值。我们发现重要性比率的自然缩放与令牌熵系统地变化。在异步动态下,这种熵比率缩放决定了两种不同的现象:在低熵时,固有的训练-推理差异会急剧放大为大量采样噪声;在高熵时,进行中的权重更新会自然地引发明显的、合法的探索性偏差。因此,仅幅度校正会无意中接受放大的噪声,同时严格掩盖由进行中的更新触发的基本探索。为了解决这个问题,我们提出了熵缩放信任区域(ESTR),它通过每个令牌的局部熵来缩放其离策略偏差,不需要辅助前向传递或显式版本切换检测。在长期代理任务和数学推理基准测试中,ESTR始终优于现有的异步方法,并实现了最佳的训练-推理一致性。与同步GRPO相比,ESTR在提高训练速度2.6倍的同时达到了可比的精度。

英文摘要

Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data destabilizes optimization and can cause policy collapse. Existing methods gate tokens by ratio magnitude alone, applying one threshold at every position. We show that the ratio's natural scale is set by token entropy, so deviations from mid-trajectory weight updates stay within this scale and carry genuine exploration. We further identify an overlooked low-entropy regime that breaks this scaling, where a near-zero probability amplifies train--inference mismatch into noise far beyond what the local entropy admits. A magnitude threshold admits this noise and discards the exploration. We therefore propose the Entropy-Normalized Trust Region (ENTR). Across long-horizon agentic tasks and mathematical reasoning benchmarks, ENTR outperforms existing asynchronous methods. It improves avg@1 on BrowseComp-Plus by $6.9\%$ over the strongest baseline, trains stably up to $30$ policy versions of staleness, and matches synchronous GRPO at a $2.6\times$ speedup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑