发表机构
Beijing Institute of Technology, Zhuhai; Dobot Robotics; Zhejiang University; Fudan University; Osaka University(北京理工大学(珠海); 越疆机器人; 浙江大学; 复旦大学; 大阪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出OPTS-TTPO,利用同策略并行树搜索增强有限样本策略梯度学习,通过分支聚合引理控制梯度偏差,在MuJoCo、Atari-57和Qwen3上显著提升性能。
AI 中文摘要
策略梯度定理给出了当前策略下的精确梯度,但有限的同策略样本可能会遗漏罕见的高回报轨迹。我们研究在固定预算内,树搜索能否提高这些轨迹的覆盖率,同时控制梯度偏差。我们引入了同策略并行树搜索(OPTS)和树轨迹策略优化(TTPO),它们使用同策略树轨迹,即在访问过的状态处从当前策略采样新的后缀。这无需动作分布校正,尽管分支会改变状态访问分布。我们的分支聚合引理表明,当分支选择和权重在输出转移采样之前固定时,分支加权树统计量可以恢复链式期望。OPTS使用估计的性能差异来选择扩展状态。在确定性动力学、精确值和最大备份优势下,诱导搜索策略的期望回报随预算单调提高。我们界定了自适应扩展带来的梯度偏差,并表明最大备份将前缀信用分配给导致更好发现后缀的动作。与有限链参考相比,TTPG的测量偏差保持在其无分支水平附近,而NaivePG的偏差从0.1251增长到0.4884。在匹配预算下,奖励引导和值引导的OPTS相比独立采样提高了正确答案覆盖率和多数投票准确率。在匹配分支数下,OPTS+TTPG相比固定分支+TTPG以适度偏差增加获得覆盖率提升。在匹配交互或 rollout 预算下,OPTS-TTPO在MuJoCo尾部回报上相比PPO提升高达28.6%,在Atari-57上以最后100日志平均回报指标取得34胜22负1平的战绩,并在所有四个Qwen3模型上相比PPO提高了微平均avg@32和pass@32。
英文摘要
The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We study whether tree search improves their coverage within a fixed budget while controlling gradient bias. We introduce On-Policy Parallel Tree Search (OPTS) and Tree Trajectory Policy Optimization (TTPO) using on-policy tree trajectories, which sample new suffixes from the current policy at visited states. This needs no action-distribution correction, although branching changes state visitation. Our Branch Aggregation Lemma shows that branch-weighted tree statistics recover chain expectations when branch choices and weights are fixed before outgoing transitions are sampled. OPTS selects expansion states using estimated performance differences. Under deterministic dynamics, exact values, and max-backup advantages, the induced search policy's expected return improves monotonically with the budget. We bound the gradient bias from adaptive expansion and show that max backup assigns prefix credit to actions leading to better discovered suffixes. Against a finite chain reference, TTPG's measured bias stays near its no-branching level, while NaivePG's bias grows from 0.1251 to 0.4884. At matched budgets, reward- and value-guided OPTS improve correct-answer coverage and majority-vote accuracy over independent sampling. At matched branch counts, OPTS + TTPG gains coverage with a modest bias increase relative to Fixed-branch + TTPG. Under matched interaction or rollout budgets, OPTS-TTPO improves MuJoCo tail returns over PPO by up to 28.6%, achieves a 34-22-1 win-loss-tie record against PPO on Atari-57 under the last-100-log mean-return metric, and improves micro-averaged avg@32 and pass@32 over PPO across all four Qwen3 models.
Comments42 pages, 12 figures