arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29961cs.LGeess.SPmath.OC

随机算子带自举的收缩框架:应用于时序差分学习

A Contraction Framework for Stochastic Operators with Bootstrapping: Application to TD Learning

Ids van der Werf, Sergio Rozada, Antonio G. Marques

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一个通用收缩框架,将带自举的随机更新建模为随机算子,在无需梯度结构且允许采样误差增长的情况下,推导出TD学习等算法的有限时间收敛界,并统一了现有确定性及随机梯度型结果。

中文摘要 AI 辅助

许多迭代算法依赖于自举(bootstrapping)。一个变量使用第二个冻结的副本作为目标进行更新,该副本会定期被更新后的变量替换。主化-最小化(majorize-minimize)方法和不精确近端点方法共享这种结构,时序差分(TD)学习也是如此。然而,对于将采样更新与每$K$步才刷新目标相结合的场景,现有的收敛性保证依赖于更新的特定结构,例如线性近似或基于梯度的内步,以及均匀有界的采样误差。我们转而将采样更新建模为参数空间上的随机算子,这将分析简化为一个不需要梯度结构的收缩论证,并允许采样误差随迭代次数增长。在此框架内,我们推导了独立同分布样本和任意目标更新周期$K$的有限时间界。我们证明,只要对冻结目标的敏感性小于内映射的收缩余量,迭代在均方根意义上几何收敛到不动点周围的球。现有的确定性冻结目标收缩和随机梯度型界作为我们框架的特例,并且TD学习的模拟重现了预测的收缩率和误差下限随步长的缩放。

英文摘要

Many iterative algorithms rely on bootstrapping. A variable is updated using a second, frozen copy as a target, which is periodically replaced with the updated variable. Majorize-minimize and inexact proximal-point methods share this structure, as does temporal-difference (TD) learning. However, existing convergence guarantees for scenarios that combine sampled updates with targets refreshed only every $K$ steps rely on the specific structure of the update, such as linear approximation or gradient-based inner steps, and on uniformly bounded sampling error. We instead model the sampled update as a stochastic operator on the parameter space, which reduces the analysis to a contraction argument that needs no gradient structure and allows the sampling error to grow with the iterates. Within this framework, we derive a finite-time bound for i.i.d. samples and any target-update period $K$. We show that the iterates converge geometrically in root mean square to a ball around the fixed point, provided the sensitivity to the frozen target is smaller than the contraction slack of the inner map. Existing deterministic frozen-target contraction and stochastic-gradient-type bounds follow as special cases of our framework, and simulations of TD learning reproduce the predicted contraction rate and scaling of the error floor with the step size.

发表机构

  • Delft University of Technology(代尔夫特理工大学)
  • Rey Juan Carlos University(胡安卡洛斯国王大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑