arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

异步TD学习在马尔可夫数据下的尖锐统计速率

Sharp Statistical Rates for Asynchronous TD Learning with Markovian Data

Yang Peng

arXiv 2609.38880首次发表:更新:

发表机构

Yau Mathematical Sciences Center, Tsinghua University(清华大学丘成桐数学科学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文证明异步TD学习在马尔可夫数据下最后一次迭代的尖锐统计速率,包含三次有效视界和混合瞬态项,并给出匹配的极小极大下界。

AI 中文摘要

我们研究从有限马尔可夫奖励过程的单条轨迹中,标准表格时间差分(TD)学习的最后一次迭代。对于折扣因子 $\gamma$,记 $H=(1-\gamma)^{-1}$,并令 $\mu_{\min}$ 和 $t_{\operatorname{mix}}$ 分别表示最小平稳概率和全变差混合时间。我们证明,对于 $0<\varepsilon\leq1$,最后一次迭代的TD以高概率实现上确界范数误差至多 $\varepsilon$,所需转移次数为 $\widetilde O\left( \frac{H^3}{\mu_{\min}\varepsilon^2} +\frac{t_{\operatorname{mix}}}{\mu_{\min}} \right)$。该速率既适用于为目标精度选择的常数步长,也适用于与目标精度和终止时间无关的递减调度。后者在显式瞬态阈值之后的所有时间上给出同时保证。统计项保留了同步TD的三次有效视界依赖性,加性混合瞬态没有额外的视界因子。该结果允许不可逆链、任意初始状态分布以及可能依赖于下一状态的有界奖励。证明使用反向时间中的锚定局部泊松方程来控制随机波动,而无需混合时间因子,并使用命中时间补偿恒等式来界定初始化误差。后者还以最坏期望反向命中时间的形式产生更精细的瞬态。对期望累积传播质量的一个界限将此论证扩展到递减步长。一个具有已知确定性奖励的三状态构造,在慢混合参数机制下,针对指定模型类,给出了统计项和混合项(对数因子内)的匹配极小极大下界。

英文摘要

We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discount factor $γ$, write $H=(1-γ)^{-1}$, and let $μ_{\min}$ and $t_{\operatorname{mix}}$ denote the minimum stationary probability and total-variation mixing time. We prove that last-iterate TD achieves sup-norm error at most $\varepsilon$ with high probability using$\widetilde O\left( \frac{H^3}{μ_{\min}\varepsilon^2} +\frac{t_{\operatorname{mix}}}{μ_{\min}} \right)$ transitions, for $0<\varepsilon\leq1$. This rate holds both for a constant step size selected for the target accuracy and for a decreasing schedule independent of the target accuracy and terminal time. The latter gives a simultaneous guarantee over all times beyond an explicit transient threshold. The statistical term retains the cubic effective-horizon dependence of synchronous TD, and the additive mixing transient has no extra horizon factor. The result allows non-reversible chains, arbitrary initial state distributions, and bounded rewards that may depend on the next state. The proof uses an anchored local Poisson equation in reverse time to control stochastic fluctuations without a mixing-time factor, and a hitting-time compensation identity to bound initialization error. The latter also yields a finer transient in terms of the worst expected reverse hitting time. A bound on the expected cumulative propagation mass extends this argument to decreasing step sizes. A three-state construction with known deterministic rewards gives matching minimax lower bounds for the statistical and mixing terms, up to logarithms, over specified model classes in a slow-mixing parameter regime.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑