arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16182cs.LGcs.AI

通过受控自举和调节值动态理解并稳定深度Q学习

Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated Value Dynamics

  • School of Computer Science, Peking University(北京大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

Bozhou Chen, Yongyi Wang, Hanyu Liu, Xionghui Yang, Wenxin Li

AI总结:

本研究系统分析深度Q学习的不稳定性来源,提出受控自举等稳定原则,在Atari-100K和Procgen上实现了竞争力性能与更优训练稳定性。

AI中文摘要:

深度Q学习(DQL)在强化学习中取得了显著的经验成功,但其训练过程却以不稳定著称。现有研究常将不稳定性归因于高估偏差或表示学习问题等孤立因素,缺乏对递归值估计过程中不同不稳定来源如何相互作用的统一理解。本研究从三个互补视角对深度Q学习的不稳定性进行系统分析:贝尔曼自举的算子级偏差、贪婪动作选择对回归噪声的估计器级敏感性,以及在激进数据复用下的参数动态失衡。我们识别出一种奖励触发的自我强化陷阱和特征性参数尖峰动态,进而推导了受控自举、分位数集成估计和基于尖峰的参数调节的稳定原则。在Atari-100K和Procgen上的实验表明,该方法具有竞争力的性能和更优的训练稳定性。

英文摘要:

Deep Q-learning (DQL) has achieved remarkable empirical success in reinforcement learning, yet its training process remains notoriously unstable. Existing studies often attribute instability to isolated factors such as overestimation bias or representation learning issues, lacking a unified understanding of how different sources of instability interact during recursive value estimation. In this work, we provide a systematic analysis of instability in deep Q-learning from three complementary perspectives: operator-level bias in Bellman bootstrapping, estimator-level sensitivity of greedy action selection to regression noise, and parameter-dynamics imbalance under aggressive data reuse. We identify a reward-triggered self-reinforcing trap and characteristic parameter spike dynamics, then derive stabilization principles for controlled bootstrapping, ensemble quantile estimation, and spike-based parameter regulation. Experiments on Atari-100K and Procgen demonstrate competitive performance and improved training stability.

↑