分布时序差分学习中的在线推断
Online Inference in Distributional Temporal-Difference Learning
浏览论文内容
中文总结 AI 辅助
该研究针对固定策略下回报分布泛函的在线推断,通过非参数分布时序差分学习估计回报分布,证明了Polyak–Ruppert平均估计量的收敛性,为光滑泛函提供自举推断依据,并建立非光滑泛函的局部渐近理论以开展推断。
中文摘要 AI 辅助
我们研究固定策略下回报分布泛函的在线统计推断问题。通过从单一马尔可夫轨迹中进行非参数分布时序差分学习来估计回报分布。针对Polyak–Ruppert平均估计量,我们证明其T次方根误差在Cramér空间中弱收敛于中心化高斯随机元素;还证明在给定观测轨迹的条件下,自举估计量与原始平均的T次方根差弱收敛于同一高斯极限。这些结果为方差、CVaR、期望短缺和期望分位数等光滑统计泛函的自举推断提供了理论依据。对于非光滑统计泛函,我们针对有限多个阈值的T^{-1/2}邻域内的估计回报CDF,建立了局部渐近理论及其自举对应形式,该理论可用于对由CDF方程表征的非光滑统计泛函(包括回报分位数)进行推断。
英文摘要
We study online statistical inference for functionals of the return distribution under a fixed policy. The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Markov trajectory. For the Polyak--Ruppert averaged estimator, we prove that its root-$T$ error converges weakly to a centered Gaussian random element in Cramér space. We also prove that, conditionally on the observed trajectory, the root-$T$ difference between the bootstrap and original averages converges weakly to the same Gaussian limit. These results justify bootstrap inference for smooth statistical functionals, including variance, CVaR, expected shortfall, and expectiles. For nonsmooth statistical functionals, we develop a local asymptotic theory for the estimated return CDF over $T^{-1/2}$-neighborhoods of finitely many thresholds, together with its bootstrap analogue. This theory allows us to conduct inference for nonsmooth statistical functionals characterized by CDF equations, including return quantiles.
发表机构
- Yau Mathematical Sciences Center, Tsinghua University(清华大学丘成桐数学科学中心)
- School of Statistics and Data Science, Shanghai University of Finance and Economics(上海财经大学统计与数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。