arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

门控Q学习:为Q学习添加离策略偏差

Gated Q-learning: Add Off-Policy Bias to Taste

Brett Daley

arXiv 2607.28916首次发表:更新:

AI 中文总结

该研究针对Q学习中离策略偏差与多步信用分配的张力,提出门控Q学习框架,通过状态-动作依赖的门控机制替代重要性采样,实现平滑插值,提升初始学习速度并支持定制化视界与偏差量。

AI 中文摘要

多步信用分配对于样本高效的强化学习至关重要,但管理Q学习中的离策略偏差仍是一个基础性挑战。30年来,从业者只能在两种选择中二选一:消除偏差但导致资格迹严重截断(Watkins的Q(λ)),或忽略偏差以更快学习但向价值估计中注入有害误差(Peng的Q(λ))。现代离策略估计器无法解决这种张力,因为重要性采样比在Q学习的贪心目标策略下会崩溃。我们引入门控Q学习(Gated Q-learning),这是一种新颖的算法框架,通过在两种历史极端情况之间平滑插值来终结这一困境。我们的方法不依赖重要性采样,而是采用一种连续的、依赖状态-动作的门控机制,以感知探索的方式选择性衰减资格迹。我们为该机制提供了严格的理论基础,证明期望算子仍是压缩映射并推导其精确不动点。实证评估验证,中间门控安全地实现了更长的信用分配视界,相比两种极端情况能更快启动学习。门控Q学习为重要性采样提供了一种简单替代方案,同时支持定制Q学习智能体的有效多步视界和离策略偏差量。

英文摘要

Multistep credit assignment is critical for sample-efficient reinforcement learning, yet managing off-policy bias in Q-learning remains a fundamental challenge. For 30 years, practitioners have been limited to a binary choice: eliminate the bias at the cost of severely truncated eligibility traces (Watkins' Q($λ$)), or ignore the bias to learn faster while injecting detrimental errors into the value estimates (Peng's Q($λ$)). Modern off-policy estimators fail to resolve this tension, as importance-sampling ratios collapse under Q-learning's greedy target policy. We introduce Gated Q-learning, a novel algorithmic framework that ends this dilemma by smoothly interpolating between the two historical extremes. Rather than relying on importance sampling, our approach employs a continuous, state-action-dependent gating mechanism to selectively attenuate eligibility traces in an exploration-aware manner. We provide a rigorous theoretical foundation for this mechanism, proving that the expected operator remains a contraction mapping and deriving its exact fixed point. Empirical evaluations verify that intermediate gating safely enables longer credit-assignment horizons, yielding faster initial learning than either extreme. Gated Q-learning offers a simple alternative to importance sampling while enabling customization of the effective multistep horizon and the amount of off-policy bias in Q-learning agents.

Comments18 pages, 2 figures, 2 tables. Published in the Reinforcement Learning Journal (RLJ); presented at the Reinforcement Learning Conference (RLC 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑