arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

极限核 Q($λ$):桥接短视界与长视界

Limiting-Kernel Q($λ$): Bridging Short and Long Horizons

Tolga Ok, Arman Sharifi Kolarijani, Peyman Mohajerin Esfahani, Mohamad Amin Sharifi Kolarijani

arXiv 2609.27741首次发表:更新:

发表机构

Delft University of Technology; Alpha Brain Technologies; University of Toronto(代尔夫特理工大学; Alpha Brain Technologies; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出极限核 Q($λ$)(LKQL)离策略价值估计器,结合 $n$ 步截断与极限核长视界近似,在保持计算复杂度的同时加速策略评估,并在 MuJoCo 基准上验证其优于 $n$ 步基线,尤其擅长长视界任务。

AI 中文摘要

在基于值的强化学习中,提高策略评估的准确性已被证明能改善下游策略优化的性能。广泛采用的基于 $n$ 步截断的近似族能产生计算高效的价值估计器,但本质上局限于较短的评估视界。相比之下,利用转移动力学全局结构的方法可以加速策略评估,但其内存和计算需求往往限制了在大型或连续状态空间上的可扩展性。为调和这些局限,我们引入了极限核 Q($λ$)(LKQL),一种离策略价值估计器,它结合了 $n$ 步截断与基于极限核(LK)的长视界近似。LKQL 具有与 $n$ 步估计器相同的复杂度阶,并可直接集成到在线和离线的 actor-critic 算法中。我们证明,在非周期性和近在线策略条件下,对于足够大的 $n$,LKQL 所基于的算子比其截断对应物提高了策略评估的收敛速度,并且在固定行为策略下,LKQL 本身在有限马尔可夫决策过程(MDPs)中几乎必然收敛到最优值。在 MuJoCo 连续控制基准上,我们展示了 LKQL 在大多数设置中优于 $n$ 步基线,尤其是在长视界任务上。

英文摘要

In value-based reinforcement learning, improving the accuracy of policy evaluation has been shown to improve downstream policy optimization performance. The widely adopted family of approximations relying on $n$-step truncation yields computationally efficient value estimators but is inherently limited to a short evaluation horizon. In contrast, methods that exploit the global structure of the transition dynamics can accelerate policy evaluation, but their memory and computational requirements often limit scalability to large or continuous state spaces. To reconcile these limitations, we introduce Limiting-Kernel Q($λ$) (LKQL), an off-policy value estimator that combines $n$-step truncation with a long-horizon approximation based on the limiting kernel (LK). LKQL has the same order of complexity as $n$-step estimators and integrates directly into both on- and off-policy actor-critic algorithms. We prove that, under aperiodicity and in the near-on-policy regime, the operator underlying LKQL improves the policy evaluation convergence rate over its truncated counterpart for sufficiently large $n$, and that LKQL itself converges almost surely to the optimal values in finite Markov decision processes (MDPs) under a fixed behavior policy. On the MuJoCo continuous-control benchmark, we show that LKQL improves over $n$-step baselines in most settings, particularly on long-horizon tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑