发表机构
Chongqing Institute of Green and Intelligent Technology; University of Science and Technology of China(重庆绿色智能技术研究院; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对大动作空间下Q学习的高估偏差问题,提出动作交集策略实现半解耦,在表格型与深度强化学习实验中大幅优于多个SOTA基线方法。
AI 中文摘要
本文针对大动作空间场景下Q学习的高估偏差问题展开研究,旨在缓解现有方法的瓶颈。研究发现,大动作空间会增大Q值估计的随机性,该随机性使得现有两类主导高估偏差研究的范式各自存在瓶颈:耦合范式(最优动作及其Q值由同一Q函数估计)始终存在正偏差,原因是随机性导致部分动作的估计值异常高于真实值,耦合方法会偏向这些动作;解耦范式(最优动作及其Q值由两个独立的Q函数估计)始终存在负偏差,原因是随机性增大了同一动作对应的两个独立Q表的估计差距。本文提出动作交集作为一种简单且有效的策略来缓解上述瓶颈,该策略通过两种设计实现半解耦:(1)允许两个Q函数共享一定比例的轨迹数据;(2)若数据样本被共享,每个Q函数采用耦合范式更新,否则采用解耦范式更新。动作交集策略的两大特性使其性能优异:(1)偏差范围大,即通过调整数据共享比例,估计偏差可从低估变化到高估;(2)粒度精细,即动作交集的大小可被任意细化,以实现更精细的控制。本文考虑了表格型强化学习和深度强化学习两种实验设置,深度强化学习实验表明,所提方法大幅优于多个SOTA基线方法;表格型实验则揭示了该方法为何能取得更优性能。
英文摘要
This paper considers the overestimation bias problem of Q-learning in the setting of a large action space, for the purpose of relieving the bottleneck of existing methods. We find that the large action space increases the randomness in Q-value estimation. The randomness makes two paradigms that drive the major literature on the overestimation problem have their own bottlenecks: the coupling paradigm, i.e., the optimal action and its Q-value are estimated with the same Q-function, always has a positive bias. This is because randomness leads to some actions having abnormally high estimated values than their true values, and the coupling methods prefer these actions. The decoupling paradigm, i.e., the optimal action and its Q-value are estimated with two independent Q-functions, always has a negative bias. This is because randomness increases the estimation gap between the two independent Q-tables for the same action. This paper shows that action intersection can be a simple yet powerful strategy to relieve these bottlenecks. The action intersection strategy enables semi-decoupling via two designs: (1) it allows two Q-functions to share a certain fraction of trajectory data; (2) if a data sample is shared, each Q-function is updated using the coupling paradigm; otherwise, using the decoupling paradigm. Two properties make the action intersection strategy powerful: (1) attaining a large bias range, i.e., varying the data sharing fraction, the estimation bias varies from underestimating to overestimating; (2) fine granularity: the action intersection size can be made arbitrarily finer to enable finer control. We consider two experiment settings, i.e., tabular and deep RL, deep RL experiments show that our method outperforms several SOTA baselines drastically; tabular experiments reveal why our method can achieve superior performance.