arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10896stat.MLcs.LG

马尔可夫采样下常步长时间差分学习的自归一化推断

Self-Normalized Inference for Constant-Stepsize Temporal-Difference Learning under Markovian Sampling

Min Zeng, Yichen Zhang, Xiaofeng Shao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对马尔可夫采样下常步长时间差分学习,提出无需估计长期协方差的自归一化推断方法,通过并行RR递归实现,在FrozenLake等数据集上验证了其有效性。

中文摘要 AI 辅助

常步长时间差分(TD)学习适用于策略评估,但从单一马尔可夫轨迹进行推断时必须考虑序列相关性和依赖步长的平稳目标。对于固定步长线性TD,我们建立了一个函数中心极限定理,其协方差保留了由随机TD矩阵和平稳迭代误差诱导的乘性分量。随后,我们推导了由同一轨迹驱动的并行理查森-龙贝格(RR)递归的联合函数极限。布朗桥自归一化器可对预先指定的状态值对比产生渐近 pivotal 置信区域,无需估计长期协方差或选择带宽、批长度。对于此类对比,该方法支持单遍实现,其内存不随轨迹长度增长。在固定步长下,推断中心为RR平稳目标。我们还研究了按时间步索引的设计,其中步长在每次运行内保持恒定,在更长时间步内减小。在显式的RR依赖速率窗口下,残差RR目标偏移、乘性余项和初始化效应在根n尺度上可忽略,从而可对投影贝尔曼解进行推断。在FrozenLake和Garnet上的实验展示了平稳目标覆盖率、RR目标校正以及按时间步索引设计的有限样本行为。

英文摘要

Constant-stepsize temporal-difference (TD) learning is attractive for policy evaluation, but inference from a single Markov trajectory must account for serial dependence and a stepsize-dependent stationary target. For fixed-stepsize linear TD, we establish a functional central limit theorem whose covariance retains the multiplicative component induced by the random TD matrix and the stationary iterate error. We then derive a joint functional limit for parallel Richardson--Romberg (RR) recursions driven by the same trajectory. A Brownian-bridge self-normalizer yields asymptotically pivotal confidence regions for prespecified state-value contrasts without estimating the long-run covariance or selecting a bandwidth or batch length. For such a contrast, the procedure admits a one-pass implementation whose memory does not grow with the trajectory length. At a fixed stepsize, the inferential center is the RR stationary target. We also study horizon-indexed designs in which the stepsize remains constant within each run and decreases across longer horizons. Under an explicit RR-dependent rate window, the residual RR target shift, multiplicative remainder, and initialization effect are negligible at the root-$n$ scale, yielding inference for the projected Bellman solution. Experiments on FrozenLake and Garnet illustrate stationary-target coverage, RR target correction, and the finite-sample behavior of the horizon-indexed design.

发表机构

  • City University of Hong Kong(香港城市大学)
  • Purdue University(普渡大学)
  • Washington University in St. Louis(圣路易斯华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑