发表机构
Peking University; Tsinghua University(北京大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于分位数的分布强化学习的统计效率,通过分位数投影分布贝尔曼方程构建估计器,建立非渐近误差界,推导渐近分布,刻画半参数效率界,研究分位数数量发散情况并建立定理,为相关推断提供基础。
AI 中文摘要
本文从统计效率角度研究基于分位数的分布强化学习。聚焦分布策略评估,目标是刻画回报分布。通过分位数投影分布贝尔曼方程的分位数不动点\(\eta_m\)获得回报分布的有限维表示,假设可访问生成模型,基于经验马尔可夫决策过程构建估计器\(\eta_m^{(n)}\)。在固定分位数数量\(m\)时,建立了\(\eta_m^{(n)}\)和\(\eta_m\)在\(W_\infty\)度量下的非渐近误差界,表明估计误差与\(m\)和\(n\)的关系为\(\widetilde{O}(\sqrt{m/n})\)。推导了分位数参数\(\sqrt{n}(\theta_m^{(n)} - \theta_m)\)的渐近分布并刻画半参数效率界。还研究了分位数数量发散的渐近情况,建立了关于平滑泛函\(\sqrt{n}(\eta_m^{(n)}(s) - \eta_m(s))f\)的Berry - Esseen定理,为分位数投影回报分布泛函的统计有效推断提供基础。
英文摘要
In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goal is to characterize the return distribution, namely the distribution of discounted cumulative rewards under a given policy. To obtain a finite-dimensional representation of the return distribution, we consider the quantile fixed point $η_m$ induced by the quantile-projected distributional Bellman equation. Assuming access to a generative model, we construct an estimator $η_m^{(n)}$ based on an empirical Markov decision process. For a fixed number of quantiles $m$, we establish a non-asymptotic error bound for $η_m^{(n)}$ and $η_m$ under the supremum $W_\infty$ metric, showing that the estimation error scales as $\widetilde{O}(\sqrt{m/n})$ with respect to $m$ and $n$. This implies that the quantile-based distributional policy evaluation problem can be solved with sample efficiency, achieving the optimal parametric $\sqrt{n}$ convergence rate. We derive the asymptotic distribution of the quantile parameters $\sqrt{n}(θ_m^{(n)}-θ_m)$ and characterize the semiparametric efficiency bound, which is attained by our estimator. Beyond the fixed-dimensional setting, we investigate the asymptotic regime in which the number of quantiles diverges. We characterize the limit covariance structure and show that it matches the semiparametric efficiency bound of the nonparametric model for distributional policy evaluation, showing that quantile-based estimators remain asymptotically efficient in the infinite-dimensional limit. Finally, we establish a Berry--Esseen theorem for smooth functionals $\sqrt{n}(η_m^{(n)}(s)-η_m(s))f$, thereby providing a foundation for statistically valid inference on functionals of the quantile-projected return distribution.