发表机构
School of Mathematical Sciences, Peking University; University of Pennsylvania; Yau Mathematical Sciences Center, Tsinghua University(北京大学数学科学学院; 宾夕法尼亚大学; 清华大学丘成桐数学科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对表格型分布强化学习的QTD建立全局有限样本保证,通过分离两种稳定性机制推导波动阶,明确区分局部随机波动与全局样本复杂度。
AI 中文摘要
我们针对表格型分布强化学习中的同步分位数时间差分学习(QTD)建立了全局有限样本保证。证明将两种稳定性机制分开:基于奖励累积分布函数的序单调性和分布Bellman算子的W∞压缩性的全局比较论证,可将任意初始化的迭代带入局部邻域;在该邻域内,我们将QTD平均场线性化,其雅可比矩阵为非奇异M矩阵,相关正半群允许进行方差敏感的鞅分析。对于步长α_t=c(t+1)^{-a}(其中a∈(1/2,1)),最终迭代的主导波动阶为Õ(T^{-a/2}/√(1−γ)),且不依赖分位数数量的多项式;确定性瞬态和所需的预热期仍可依赖最小Bellman目标密度,最坏情况下其阶为m^{-1}。因此,该结果明确区分了局部随机波动与全局样本复杂度。
英文摘要
Quantile temporal-difference learning (QTD) is an effective method for learning return distributions through quantile approximation, yet its finite-time behavior remains poorly understood. Its update is nonlinear and nonsmooth, and the stability needed for a sharp convergence rate holds only near the target. We establish a global high-probability last-iterate guarantee for synchronous tabular QTD under general positive, nonincreasing step-size sequences and arbitrary initialization in the natural parameter range. For polynomially decaying step sizes with exponent $a\in(0,1)$, the last iterate converges to the target at rate $T^{-a/2}$ in the infinity norm, up to logarithmic and lower-order terms. A suitably tuned harmonic schedule recovers the $T^{-1/2}$ statistical rate up to logarithmic factors. For the $m$-quantile representation, its $\infty$-Wasserstein error scales as $\sqrt{m/T}$ up to logarithmic factors, matching the leading polynomial dependence on the quantile resolution and sample size of the corresponding model-based estimator. The proof uses a two-stage global-to-local argument. From arbitrary initialization, Bellman contraction and CDF monotonicity first bring the iterate close to the target, after which, a novel variance--drift matching argument sharpens the control of accumulated noise and local contraction reduces the remaining errors, yielding the sharp rate. Simulations verify the predicted polynomial decay and assess the finite-time entrance bound.