带马尔可夫噪声的投影双时间尺度随机逼近的有限时间集中性与收敛速率
Finite-Time Concentration and Convergence Rates for Projected Two-Time-Scale Stochastic Approximation with Markov Noise
- Laboratoire de Recherche de l’EPITA(EPITA研究实验室)
- Indian Institute of Technology Bombay(印度理工学院孟买分校)
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对带马尔可夫噪声的投影双时间尺度随机逼近,利用Skorokhod映射建立有限时间高概率界,证明几乎必然收敛并给出显式速率,应用于演员-评论家等算法。
AI中文摘要:
我们研究由受控马尔可夫链驱动的投影双时间尺度随机逼近的有限时间集中性与收敛速率。平均化的快映射是压缩的,而慢迭代被投影到紧凸多面体上。相应的投影常微分方程在边界处可能具有不连续向量场,这阻碍了基于Lipschitz向量场的标准分析的直接应用。利用Skorokhod映射,我们建立了跟踪移动快平衡和投影慢动力学的显式高概率界。这些界将鞅波动、马尔可夫噪声残差和时间尺度分离引起的偏差分开。满足在固定时间区间上均匀下降条件的Lipschitz Lyapunov函数可得出几乎必然收敛,当下降具有幂次下界时,给出显式的最后迭代速率。在均匀Lyapunov压缩下,多项式步长产生联合快跟踪和慢Lyapunov误差指数,任意接近1/3。在额外假设约化慢更新映射是欧几里得压缩的情况下,对数分离的步长将联合速率提高到几乎必然的O(n^{-1/2}\log n),包括边界平衡点。同样的速率在涉及盒子的严格吸引面和在该面上恒定的快平衡的不同几何条件下也成立。演员-评论家应用在约束策略类内相对于最优值实现了几乎必然的值差距速率O(n^{-1}\log n)。进一步的应用包括投影TD(0)和投影随机梯度下降。我们还将分析扩展到欧几里得压缩性下快递归的投影。
英文摘要:
We study finite-time concentration and convergence rates for projected two-time-scale stochastic approximation driven by a controlled Markov chain. The averaged fast map is contractive, while the slow iterate is projected onto a compact convex polyhedron. The associated projected ordinary differential equation may have a discontinuous vector field at the boundary, preventing a direct application of standard analyses based on Lipschitz vector fields. Using the Skorokhod map, we establish explicit high-probability bounds for tracking the moving fast equilibrium and the projected slow dynamics. These bounds separate martingale fluctuations, Markov-noise residuals, and the bias due to time-scale separation. A Lipschitz Lyapunov function satisfying a uniform decrease condition over fixed time intervals yields almost-sure convergence, with explicit last-iterate rates when the decrease admits a power lower bound. Under uniform Lyapunov contraction, polynomial step sizes yield joint fast-tracking and slow Lyapunov-error exponents arbitrarily close to $1/3$. Under the additional assumption that the reduced slow update map is a Euclidean contraction, logarithmically separated step sizes improve the joint rate to $O(n^{-1/2}\log n)$ almost surely, including for boundary equilibria. The same rate holds under a distinct geometric condition involving a strictly attracting face of a box and a fast equilibrium that is constant on that face. An actor-critic application achieves an almost-sure value-gap rate of $O(n^{-1}\log n)$ relative to the optimum within the constrained policy class. Further applications include projected TD(0) and projected stochastic gradient descent. We also extend the analysis to projection of the fast recursion under Euclidean contractivity.