arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在Q学习中不确定性下的TD目标聚合再探讨

Revisiting TD Target Aggregation under Uncertainty in Q-Learning

Lipeng Zu, Xiaonan Zhang

arXiv 2608.03069首次发表:更新:

AI 中文总结

本研究针对Q学习中TD目标聚合受估计噪声放大误差的问题,提出SADQ方法,通过动力学模型的单步展开预测正则化TD目标,在多类任务上提升了DQN的训练稳定性。

AI 中文摘要

深度Q网络(DQNs)通过自举时序差分(TD)更新学习价值函数,其中未来回报通过对下一状态动作值的贪婪最大化进行近似。尽管该方法有效,但这种聚合规则本质上对估计噪声敏感:当Q值存在不确定性时,最大化算子会确定性地倾向于最大的估计值,而不考虑其可靠性,通过自举过程放大误差。本研究提出了后继展开聚合深度Q网络(SADQ),这是对Q学习的简单修改,用于正则化TD目标的形成方式。SADQ使用从学习到的动力学模型得到的单步展开预测来指导候选下一状态动作之间的比较,在不改变基础学习框架的情况下,为聚合步骤引入了额外的结构。由此产生的混合贝尔曼更新在保留标准不动点的同时,衰减了不可靠的最大值(当模型误差减小时)。我们提供的理论分析表明,SADQ以逐点方式减少自举导致的过估计。实验结果显示,与强大的DQN变体相比,SADQ在经典控制任务、基于真实世界向量的环境以及Atari基准测试中,始终提升了训练稳定性。

英文摘要

Deep Q-Networks (DQNs) learn value functions through bootstrapped temporal-difference updates, where future returns are approximated using a greedy maximization over next-state action values. While effective, this aggregation rule is inherently sensitive to estimation noise: when Q-values are uncertain, the maximization operator deterministically favors the largest estimate, regardless of its reliability, leading to amplified errors through bootstrapping. In this work, we propose the \textbf{S}uccessor Rollout \textbf{A}ggregation \textbf{D}eep \textbf{Q}-Network (SADQ), a simple modification to Q-learning that regularizes how the TD target is formed. SADQ uses one-step rollout predictions from a learned dynamics model to guide the comparison among candidate next-state actions, introducing additional structure into the aggregation step without altering the underlying learning framework. The resulting mixed Bellman update attenuates unreliable maxima while preserving the standard fixed point under diminishing model error. We provide theoretical analysis showing that SADQ reduces bootstrap-induced overestimation in a pointwise manner. Empirically, SADQ consistently improves training stability across classical control tasks, real-world vector-based environments, and Atari benchmarks when compared to strong DQN variants.

CommentsAccepted in the 42nd Conference on Uncertainty in Artificial Intelligence (UAI)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑