发表机构
School of Mathematical Sciences, Peking University; Yau Mathematical Sciences Center, Tsinghua University(北京大学数学科学学院; 清华大学丘成桐数学科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对分布强化学习的分位数时间差分学习,建立同步与异步QTD的函数中心极限定理,提出无需存储完整轨迹的在线推断方法,实现高效统计推断。
AI 中文摘要
本文研究分布强化学习中如何对分位数时间差分学习(Quantile Temporal Difference Learning, QTD)进行统计推断。假设可获取生成模型,我们首先为同步和异步QTD建立函数中心极限定理,表明QTD的平均迭代弱收敛到重标布朗运动。接着我们提出在线推断方法,该方法基于随机缩放,通过利用整个QTD路径的信息构造渐近枢轴统计量,且此统计量可在线计算,无需存储QTD迭代的完整轨迹,这大幅降低了内存需求,使分布强化学习中能进行高效统计推断。
英文摘要
In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming access to a generative model, we first establish functional central limit theorems for both synchronous and asynchronous QTD, which show that the averaged iterates of QTD converge weakly to a rescaled Brownian motion. We next provide online inference methods. Based on random scaling, the inference procedure constructs an asymptotically pivotal statistic for inference by using the information along the whole QTD path. Meanwhile, the proposed statistic can be computed online without storing the entire trajectory of QTD iterates. This substantially reduces the memory requirement and enables efficient statistical inference in distributional reinforcement learning.