发表机构
Department of Electrical Engineering, Stanford University(电气工程系,斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究乘性噪声模型下任意范数压缩映射的随机逼近,提出统一基本分析方法,通过平均噪声序列等得到范数误差的李雅普诺夫漂移不等式,进而得出均方和集中性界,还讨论了证明技术的可推广性。
AI 中文摘要
我们在乘性噪声模型下,为具有任意范数压缩映射的随机逼近(SA)建立了均方和集中性界。该噪声模型中噪声可能与迭代的范数仿射缩放,且迭代可能无界。这些设置出现在强化学习中。早期工作通过广义莫罗包络构造光滑李雅普诺夫函数来处理任意范数,通过多阶段自举论证处理乘性噪声和无界迭代。我们提出一种统一的基本分析方法来得到这两个界。利用平均噪声序列和相应辅助迭代,直接得到范数误差的一步李雅普诺夫漂移不等式。对于均方界,将漂移不等式与归纳论证结合表明迭代在期望上保持有界。对于集中性界,在一系列控制迭代的“好”事件上进行概率归纳,从而应用标准的阿祖马 - 霍夫丁界。我们的方法通过允许步长对数依赖于置信水平,得到了乘性噪声下SA的首个次高斯尾极大(全程)集中性界。此外,我们还讨论了这些证明技术对其他噪声模型和迭代算法的可推广性。
英文摘要
We establish mean-square and concentration bounds for stochastic approximation (SA) with arbitrary norm contractive mappings, under a multiplicative noise model where the noise may scale affinely with the norm of the iterates, and the iterates are potentially unbounded. These settings arise in reinforcement learning, where operators are often contractive in the $\ell_\infty$ norm and the noise scales with the iterates. To address the arbitrary norm, earlier works replace the non-smooth squared norm with a smooth Lyapunov function constructed via the generalized Moreau envelope. For concentration analysis, these works handle multiplicative noise and unbounded iterates through a multi-stage bootstrapping argument that starts from a time-varying worst-case bound and iteratively refines it. We instead present a unified and elementary analysis that yields both bounds. Using an averaged noise sequence and corresponding auxiliary iterates, we obtain a one-step Lyapunov drift inequality for the normed error directly, without smoothing the norm or constructing an envelope. For the mean-square bound, we combine this drift inequality with an induction argument showing that the iterates remain bounded in expectation. For the concentration bound, we develop a probabilistic induction over a sequence of "good" events on which the iterates are controlled, allowing the standard Azuma-Hoeffding bound to be applied. Our approach yields the first sub-Gaussian tailed maximal (all-time) concentration bound for SA under multiplicative noise, by allowing the stepsize to depend logarithmically on the confidence level. Beyond the specific setting considered here, we discuss the generalizability of these proof techniques to other noise models and iterative algorithms.
CommentsSubmitted to Stochastic Processes and their Applications