arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22636cs.LGmath.OCstat.ML

基于稳定无限维线性函数近似的Q学习

Q-Learning with Stable Infinite-Dimensional Linear Function Approximation

Shengbo Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究开发了稳定无限维线性函数近似的Q学习框架,提出两种SA算法并证明其收敛界,可自动适应几何与光滑性,还展示了其在Q测度学习等场景的应用。

中文摘要 AI 辅助

采用线性函数近似的Q学习可能不稳定,因为任意近似架构未必能保持Bellman压缩性。我们针对从单一马尔可夫行为策略轨迹进行的Q学习,开发了一种稳定的无限维线性函数近似框架。学习变量是紧致潜在度量空间(L,ρ)上的系数场θ∈C(L)。该框架使用重构算子将θ映射为连续Q函数,同时使用压缩算子将Bellman更新映射回潜在坐标。两个算子的非扩张性在C(L)上诱导出一个压缩潜在Bellman映射,其具有唯一不动点θ*,该不动点的重构结果可在表示误差范围内逼近最优Q函数。我们提出两种随机近似(SA)算法,并建立了它们的上范数收敛界,其主项阶为O~(n^(-1/2))。这种无限维公式为识别决定统计难度的结构提供了强大抽象。压缩映射关于ρ的光滑性由θ*和SA迭代继承,使得均匀估计误差可通过(L,ρ)的覆盖数而非C(L)的维度来控制。值得注意的是,我们提出的SA算法对ρ的选择不敏感,因此可自动适应光滑性和几何结构。我们还通过结合线性密度近似的Q测度学习以及冻结预训练网络下的输出层神经权重训练,进一步展示了该框架的应用。

英文摘要

Q-learning with linear function approximation can be unstable because an arbitrary approximation architecture need not preserve the Bellman contraction. We develop a stable infinite-dimensional linear function approximation framework for Q-learning from a single Markovian behavior-policy trajectory. The learning variable is a coefficient field $θ\in C(\mathbb L)$ on a compact latent metric space $(\mathbb L,ρ)$. The framework uses a reconstruction operator that maps $θ$ to a continuous Q-function and a compression operator that maps Bellman updates back to latent coordinates. Nonexpansiveness of both operators induces a contractive latent Bellman map on $C(\mathbb L)$, with a unique fixed point $θ^*$ whose reconstruction approximates the optimal Q-function up to representation error. We propose two stochastic approximation (SA) algorithms and establish their sup-norm convergence bounds with a leading term of order $\widetilde O(n^{-1/2})$. The infinite-dimensional formulation provides a powerful abstraction for identifying the structures that govern statistical difficulty. Smoothness of the compression map in $ρ$ is inherited by $θ^*$ and the SA iterates, allowing uniform estimation errors to be controlled through covering numbers of $(\mathbb L,ρ)$ rather than the dimension of $C(\mathbb L)$. Remarkably, the SA algorithms we propose are agnostic to the choice of $ρ$, and thus can automatically adapt to both the smoothness and the geometry. We further illustrate the framework through Q-measure-learning with linear density approximation and output-layer neural weight training under a frozen pretrained network.

发表机构

  • University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

↑