arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14922stat.MLcs.LGmath.OCmath.PR

随机逼近的稳态收敛性

Steady-State Convergence of Stochastic Approximation

Yixuan Zhang, Qiaomin Xie

首次发表
浏览论文内容

中文总结 AI 辅助

本文为常数步长收缩随机逼近建立了统一的稳态收敛理论,涵盖马尔可夫乘性噪声及局部不可微情形,获得最优高斯近似,并应用于异步Q学习及提出偏差减少方案。

中文摘要 AI 辅助

对于常数步长随机逼近(SA),迭代在分布上收敛到一个依赖于步长 $\alpha$ 的平稳分布。稳态收敛(SSC)关注的是当 $\alpha \downarrow 0$ 时,缩放后的平稳分布的极限。现有的SSC理论要求噪声为独立同分布(i.i.d.)或加性噪声,且均值算子全局可微,并产生次优的收敛速率。我们为受马尔可夫乘性噪声驱动的常数步长收缩SA建立了一个统一的SSC理论,涵盖了局部可微和局部不可微的均值算子。一个关键的方法论贡献是多步普适性框架,该框架逐步将原始随机递归简化为易于处理的辅助动力学,同时保持其稳态极限。在不动点处的局部二次线性化下,我们获得了缩放稳态在Wasserstein-2距离下的最优速率 $O(\sqrt{\alpha})$ 的高斯近似,这进一步为原始迭代提供了有限时间的高斯近似。在局部不可微的情况下,我们建立了一个一般的SSC结果,并表明前导渐近偏差的阶数可以是 $\sqrt{\alpha}$,这与光滑情况下的 $\alpha$ 阶偏差形成对比。我们将该理论应用于马尔可夫线性SA和异步Q学习,这两者均未被先前的结果所覆盖。我们进一步提出了一种无需知道局部光滑性状态的Q学习偏差减少方案,并通过数值实验进行了验证。

英文摘要

For constant-stepsize stochastic approximation (SA), the iterates converge in distribution to a stationary law that depends on the stepsize $α.$ Steady-state convergence (SSC) concerns the limit of the scaled stationary distribution as $α\downarrow 0.$ Existing SSC theory requires i.i.d. or additive noise and global differentiability of the mean operator, and yields suboptimal rates. We develop a unified SSC theory for constant-stepsize contractive SA driven by Markovian, multiplicative noise, covering both locally differentiable and locally nondifferentiable mean operators. A key methodological contribution is a multi-step universality framework that progressively reduces the original stochastic recursion to tractable auxiliary dynamics while preserving its steady-state limit. Under local quadratic linearization at the fixed point, we obtain a Gaussian approximation of the scaled steady state at the optimal rate $O(\sqrtα)$ in Wasserstein-2 distance, which further gives finite-time Gaussian approximations for the raw iterates. In the locally nondifferentiable regime, we establish a general SSC result and show that the leading-order asymptotic bias can be of order $\sqrtα$, in contrast to the $α$-order bias in the smooth regime. We apply the theory to Markovian linear SA and asynchronous Q-learning, neither of which is covered by prior results. We further propose a bias-reduction scheme for Q-learning that requires no knowledge of the local smoothness regime, validated by numerical experiments.

发表机构

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑