arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

处处自由,树上精确:PPO 的丢弃修正以激进复用换取样本效率

Free Everywhere, Exact on Trees: PPO's Dropped Correction Buys Sample Efficiency Under Aggressive Reuse

Nima H. Siboni

arXiv 2609.39634首次发表:更新:

AI 中文总结

本文证明在历史单射动力学下可精确恢复PPO丢弃的状态访问比率,提出单参数修正族以平衡偏差与方差,并在调度任务上显示激进复用下修正预热提升样本效率。

AI 中文摘要

常见的策略改进方法,包括 TRPO、PPO 和 GRPO,是在行为策略的状态访问分布下估计策略改进,而非在改进后策略自身的分布下进行。这种替代使得目标函数可以从行为策略的轨迹中估计,但引入了随策略散度增长的偏差,因此需要信任区域或裁剪,并且无法复用远离策略的批次。我们证明,在历史单射动力学(每个状态仅由唯一历史到达)下,被丢弃的状态访问比率等于沿采样前缀的每步策略比率的乘积,这在每条轨迹上都成立,而不仅仅在期望上。因此,该比率可以从 PPO 已计算的对数概率中精确恢复。自回归生成和规范序构造优化均属于历史单射。精确修正引入了随视界增长的重要性采样方差,因此我们将其推广为一个单参数族,以 PPO($\alpha{=}0$)和完整修正($\alpha{=}1$)为端点:一个单一的偏差-方差旋钮。对未裁剪代理目标的梯度级分析识别出修正作用的两个通道以及其携带信号的三个条件;一个可枚举的测试平台确认了这些条件的预测。在困难的信用分配调度任务上,使用激进早期样本复用的短时修正预热比 PPO 和未修正的相同复用学习得更快;边际收益随任务难度增长(学习曲线 AUC 增加 $+0.02$ 到 $+0.09$),对 PPO 的早期优势与复用引起的前缀偏差一致。全程保持修正,或在裁剪已包含复用偏差时应用修正,则效果为零甚至有害。

英文摘要

Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribution rather than the improved policy's own. The substitution makes the objective estimable from the behavioral policy's rollouts but adds a bias growing with policy divergence, hence the trust region or clip, and hence no reuse of a batch far off-policy. We show that under history-injective dynamics, where each state is reached by exactly one history, the dropped state-visitation ratio equals the product of per-step policy ratios along the sampled prefix, on every trajectory and not only in expectation. The ratio is therefore restored exactly, from log-probabilities PPO already computes. Autoregressive generation and canonical-order constructive optimization are both history-injective. The exact correction pays importance-sampling variance that grows with the horizon, so we generalize it to a one-parameter family with PPO ($α{=}0$) and the full correction ($α{=}1$) as endpoints: a single bias--variance knob. A gradient-level analysis of the unclipped surrogate identifies two channels the correction acts through and three conditions under which it carries signal; an enumerable testbed confirms the conditions' predictions. On hard credit-assignment scheduling tasks, a short corrected warmup with aggressive early sample reuse learns faster than PPO and than the same reuse uncorrected; the marginal gain grows with task difficulty ($+0.02$ to $+0.09$ learning-curve AUC), and the early win over PPO tracks the prefix bias that reuse incurs. A correction held throughout, or applied where clipping already contains the reuse bias, is null to harmful.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑