AI 中文总结
本文提出EPIG-Tree,将树状回放构建视为计算分配问题,通过总方差分解推导分支分配法则,以最小化策略梯度不确定性,实验证明其在连续控制和多轮Wordle中优于GRPO和熵分支。
AI 中文摘要
基于奖励的强化学习用于语言模型,以组相对策略优化(GRPO)为代表,将整个随机轨迹压缩为单个标量奖励。这种方法简洁且可扩展,但在探索和奖励分配上效率低下:一条轨迹可能包含许多因果决策、恢复尝试和环境随机事件,然而每个词元或动作都继承了同一轨迹级别的优势值。我们将基于树的分支构建研究为策略梯度估计的计算分配问题。我们的核心主张是:分支应放置于不仅策略不确定之处,而是额外分支在单位计算下最能降低策略梯度不确定性的位置。通过对局部策略梯度随机变量进行总方差定律分解,我们推导出两条分配法则:新分支减少决策不确定性,而重复后缀回放减少延续不确定性。由此得到的EPIG-Tree评分利用已计算的回放来分配分支。它估计占用率和得分加权的价值不确定性,以及后缀法则 $n_e \propto w_e \\|\nabla_\theta \log \pi(a_e|h_e)\\| \sigma_e / \sqrt{c_e}$。实验上,EPIG在克隆状态控制中降低了梯度均方误差,在13个环境的扫描中赢得了全部9个密集连续控制环境,并近乎完美地恢复了参考梯度方向;相对于熵分支,它改善了冻结语言模型的梯度校准。在在线单轮数学任务中,树局部信用分配优于平坦GRPO,而分支放置次于词元级信用分配。在在线多轮Wordle任务中,EPIG达到了最高的最终胜率(0.850),超越了平坦GRPO(其早期饱和于0.790)以及熵分支(随着训练进行),证实了梯度估计优势可迁移至有状态、大动作空间的环境。
英文摘要
Reward-based reinforcement learning for language models, exemplified by Group Relative Policy Optimization (GRPO), collapses an entire stochastic trajectory into a single scalar reward. This is clean and scalable, but it explores and allocates reward inefficiently: a trajectory may contain many causal decisions, recovery attempts, and environment-randomness events, yet every token or action inherits one trajectory-level advantage. We study tree-based rollout construction as a compute-allocation problem for policy-gradient estimation. Our central claim is that branches should be placed not where the policy is merely uncertain, but where an additional branch most reduces uncertainty about the policy gradient per unit of compute. From a law-of-total-variance decomposition of the local policy-gradient random variable, we derive two allocation laws: new branches reduce decision uncertainty, while repeated suffix rollouts reduce continuation uncertainty. The resulting EPIG-Tree score allocates branches using the already computed rollouts. It estimates occupancy- and score-weighted value uncertainty, along with a suffix law $n_e \propto w_e \|\nabla_θ\log π(a_e|h_e)\| σ_e / \sqrt{c_e}$. Empirically, EPIG reduces gradient MSE in cloned-state control, winning in all nine dense continuous-control environments of a 13-environment sweep and recovering the reference gradient direction near-perfectly, and it improves frozen-LLM gradient calibration relative to entropy branching. In online single-turn math, tree-local credit beats flat GRPO, while branch placement is secondary to token-level credit assignment. In online multi-turn Wordle, EPIG attains the highest final win rate (0.850), overtaking flat GRPO, which saturates early at 0.790, and entropy branching as training proceeds, confirming that the gradient-estimation advantage transfers to a stateful, large-action setting.
Comments12 pages, 8 figures