arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度$V$-学习的收敛框架:误差传播与尖锐行动间隙界

A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds

Yury Kolomeytsev

arXiv 2609.18782首次发表:更新:

发表机构

Lomonosov Moscow State University(莫斯科国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文为深度$V$-学习建立收敛框架,将更新误差分解为六类残差并给出尖锐行动间隙界,证明固定水平生成式重置近似ERM过程的策略损失一致性,为FIFO/交错SGD提供残差衰减准则。

AI 中文摘要

我们建立了深度$V$-学习在水平$H$下的收敛界。该算法将标量值函数拟合到从执行转移中获得的目标上,并使用预测模型和值函数选择行动。对于当前观测到的后继目标(使用新鲜的真实核结果),条件均值为$\mathcal{T}^\beta V$,它是对行为策略行动的平均。贝尔曼最优性更新为$\mathcal{T} V$。我们将更新误差分解为六个残差:拟合、转移重用、目标构建、重放、行动选择和探索。在$L^s$可集中性下,它们的$L^p$范数($p=s/(s-1)$)控制期望的$L^1$策略损失。该界明确加权了仅来自最后$H-1$个更新块的残差,外加一个针对较短运行的初始化项。我们量化了跨水平级别共享采样分布的成本。对于阶为$n^{-\nu}$的统计误差界,我们推导出最优连续分配以及一个整数分配,其目标值在约束最优值的$2^\nu$因子内。具有指数$\alpha$的边际条件给出阶为$\Lambda^{1+\alpha/p}$的行动误差,其中$\Lambda$结合了网络漂移和评分误差;一步构造证明了该指数的尖锐性。冻结评分与最优评分之间距离的界将最优间隙条件转化为冻结迭代间隙界,同时保留最优平局的质量。部署时的生存概率和覆盖条件为使用近似评分选择的策略提供了界。每个水平级别使用独立的空间ReLU网络给出条件神经回归速率,有限状态情形给出无对数的期望拟合速率。这些结果为具有精确行动评分的固定水平生成式重置近似ERM过程提供了期望策略损失一致性,并为FIFO/交错SGD提供了显式残差衰减准则。

英文摘要

We establish convergence bounds for deep $V$-learning with horizon $H$. The algorithm fits a scalar value function to targets from executed transitions and selects actions using a predictive model and the value function. For current observed-successor targets with fresh true-kernel outcomes, the conditional mean is $\mathcal{T}^βV$, which averages over behavior-policy actions. The Bellman optimality update is $\mathcal{T} V$. We decompose the update error into six residuals: fitting, transition reuse, target construction, replay, action selection, and exploration. Under $L^s$ concentrability, their $L^p$ norms ($p=s/(s-1)$) control expected $L^1$ policy loss. The bound explicitly weights residuals from only the last $H-1$ update blocks, plus an initialization term for shorter runs. We quantify the cost of a shared sampling distribution across horizon levels. For statistical error bounds of order $n^{-ν}$, we derive optimal continuous allocations and an integer allocation whose objective is within a factor $2^ν$ of the constrained optimum. A margin condition with exponent $α$ gives action error of order $Λ^{1+α/p}$, where $Λ$ combines network drift and score error; a one-step construction proves the exponent sharp. Bounds on the distance between frozen and optimal scores transfer an optimal-gap condition to frozen-iterate gap bounds while retaining the mass of optimal ties. Survival probabilities and coverage conditions at deployment yield bounds for policies selected with approximate scores. Separate spatial ReLU networks per horizon level give a conditional neural regression rate, and the finite-state case gives a log-free expected fit rate. These results give expected policy-loss consistency for the fixed-horizon generative-reset approximate-ERM procedure with exact action scores and provide an explicit residual-decay criterion for FIFO/interleaved SGD.

Comments37 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑