AI 中文总结
针对离策略强化学习中多步回报引发的悲观偏差问题,提出ENQ算法,证明其理论性质,在27个任务上与LQL性能相当且吞吐量更高,集成多评论家时获益更多。
AI 中文摘要
多步回报可加速离策略强化学习中的奖励传播,但会将每个决策的评估与后续次优的记录动作耦合,产生随时间范围增长的悲观偏差。我们提出期望n步Q学习(ENQ),它用动作值误差上的非对称期望损失替代对称的n步时间差分(TD)损失,其中期望水平τ是除n步TD外唯一新增的方法特定超参数。我们证明ENQ算子是γⁿ收缩算子。在确定性动态下,当τ=1时,其在覆盖的支持对的最优动作值函数Q*上的偏差消失,且对应的不动点满足长程Q学习(LQL)所用的分离n实例及其下界不等式的倍数。在随机动态下,该算子偏差具有与时间范围无关的噪声常数的双边界。在27个操作和导航任务实例中,使用单一期望水平τ=0.8和固定的备份时间范围,ENQ整体与LQL相当,在我们的分析研究中实现更高的训练步骤吞吐量,且在受控规模实验中从十评论家集成中获益更多。
英文摘要
Multi-step returns accelerate reward propagation in off-policy reinforcement learning, but couple the evaluation of each decision to the suboptimal logged actions that follow it, inducing a pessimistic bias that grows with the horizon. We propose Expectile $n$-step Q-learning (ENQ), which replaces the symmetric $n$-step temporal-difference (TD) loss with an asymmetric expectile loss on the action-value error, with expectile level $τ$ as the only method-specific hyperparameter added beyond $n$-step TD. We prove that the ENQ operator is a $γ^{n}$-contraction. Under deterministic dynamics, at $τ=1$, its bias vanishes at the optimal action-value function $Q^*$ on covered in-support pairs, and the corresponding fixed point satisfies the separation-$n$ instance and its multiples of the lower-bound inequality used by Long-Horizon Q-learning (LQL). Under stochastic dynamics, the operator bias admits two-sided bounds with horizon-independent noise constants. Using a single expectile level $τ=0.8$ and a fixed backup horizon across 27 manipulation and navigation task instances, ENQ is competitive with LQL on aggregate, achieves higher measured training-step throughput in our profiling study, and benefits more from a ten-critic ensemble in a controlled scaling experiment.