发表机构
Yale University; Wuhan University(耶鲁大学; 武汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对离线RL中最优贝尔曼算子不可观测的问题,提出将算子估计与价值学习解耦的框架,用条件扩散模型估计奖励律和转移核,建立了无完备性假设的理论收敛速率,实验验证了方法有效性。
AI 中文摘要
在离线强化学习(offline RL)中,最优动作价值函数$Q^*$的估计可被表述为仅基于离线观测数据求解最优贝尔曼方程。一个核心挑战在于奖励函数和转移核是未知的,因此最优贝尔曼算子无法直接从数据中观测到。为解决该问题,我们提出了一种将算子估计与价值函数学习解耦的新框架。该方法中,我们首先构建条件扩散模型(conditional diffusion models)来估计奖励律和转移核,这会生成最优贝尔曼算子的数据驱动近似。随后,我们将这些估计器代入贝尔曼方程,通过在神经网络函数类上最小化经验贝尔曼残差,得到$Q^*$的深度估计器。理论层面,我们首先通过对条件扩散估计在总变差距离上的端到端分析,建立了学习最优贝尔曼算子的精确非渐近收敛速率;接着,针对超额贝尔曼残差风险,建立了神谕价值阶段速率$\boldsymbol{\tilde{\tmashcal O}\bigl(n^{-\frac{2\beta}{d_x+d_a+2\beta}}\bigr)}$,其中$d_x$和$d_a$分别表示状态空间和动作空间的维度,$\beta$表示$Q^*$的赫尔德光滑度指数。重要的是,我们的理论分析不依赖深度强化学习理论中常用的完备性假设。大量数值实验证明了所提方法的有效性及其出色的经验性能。
英文摘要
In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations. A fundamental challenge is that the reward function and transition kernel are unknown, so the optimal Bellman operator is not directly observable from data. To address this issue, we propose a novel framework that decouples operator estimation from value function learning. In this approach, we first formulate conditional diffusion models to estimate the reward law and transition kernel, which induces a data-driven approximation of the optimal Bellman operator. We then plug these estimators into the Bellman equation and obtain a deep estimator of $Q^*$ by minimizing the empirical Bellman residual over a neural network function class. Theoretically, we first establish sharp nonasymptotic convergence rates for learning the optimal Bellman operator through an end-to-end analysis of conditional diffusion estimation in total variation distance. We then establish the oracle value-stage rate $\widetilde{\mathcal O}\bigl(n^{-\frac{2β}{d_x+d_a+2β}}\bigr)$ for the excess Bellman residual risk. Finally, under a concentrability condition, we translate this residual bound into an $L^2$ convergence rate of $\widetilde{\mathcal O}\bigl(n^{-\fracβ{d_x+d_a+2β}}\bigr)$ for the resulting deep estimator of $Q^*$, where $d_x$ and $d_a$ denote the dimensions of the state and action spaces, respectively, and $β$ denotes the Hölder smoothness index of $Q^*$. Importantly, our theoretical analysis does not rely on completeness assumptions commonly used in deep RL theory. Extensive numerical experiments demonstrate the effectiveness of the proposed method and its strong empirical performance.