arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于扩散模型的离线深度Q*估计

Offline Deep Q* Estimation with Diffusion Models

Xiaohong Chen, Yuling Jiao, Lican Kang, Jerry Zhijian Yang, Chen Zhong

arXiv 2608.14401首次发表:更新:

发表机构

Yale University; Wuhan University(耶鲁大学; 武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对离线RL中最优贝尔曼算子不可观测的问题,提出将算子估计与价值学习解耦的框架,用条件扩散模型估计奖励律和转移核,建立了无完备性假设的理论收敛速率,实验验证了方法有效性。

AI 中文摘要

在离线强化学习(offline RL)中,最优动作价值函数$Q^*$的估计可被表述为仅基于离线观测数据求解最优贝尔曼方程。一个核心挑战在于奖励函数和转移核是未知的,因此最优贝尔曼算子无法直接从数据中观测到。为解决该问题,我们提出了一种将算子估计与价值函数学习解耦的新框架。该方法中,我们首先构建条件扩散模型(conditional diffusion models)来估计奖励律和转移核,这会生成最优贝尔曼算子的数据驱动近似。随后,我们将这些估计器代入贝尔曼方程,通过在神经网络函数类上最小化经验贝尔曼残差,得到$Q^*$的深度估计器。理论层面,我们首先通过对条件扩散估计在总变差距离上的端到端分析,建立了学习最优贝尔曼算子的精确非渐近收敛速率;接着,针对超额贝尔曼残差风险,建立了神谕价值阶段速率$\boldsymbol{\tilde{\tmashcal O}\bigl(n^{-\frac{2\beta}{d_x+d_a+2\beta}}\bigr)}$,其中$d_x$和$d_a$分别表示状态空间和动作空间的维度,$\beta$表示$Q^*$的赫尔德光滑度指数。重要的是,我们的理论分析不依赖深度强化学习理论中常用的完备性假设。大量数值实验证明了所提方法的有效性及其出色的经验性能。

英文摘要

In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations. A fundamental challenge is that the reward function and transition kernel are unknown, so the optimal Bellman operator is not directly observable from data. To address this issue, we propose a novel framework that decouples operator estimation from value function learning. In this approach, we first formulate conditional diffusion models to estimate the reward law and transition kernel, which induces a data-driven approximation of the optimal Bellman operator. We then plug these estimators into the Bellman equation and obtain a deep estimator of $Q^*$ by minimizing the empirical Bellman residual over a neural network function class. Theoretically, we first establish sharp nonasymptotic convergence rates for learning the optimal Bellman operator through an end-to-end analysis of conditional diffusion estimation in total variation distance. We then establish the oracle value-stage rate $\widetilde{\mathcal O}\bigl(n^{-\frac{2β}{d_x+d_a+2β}}\bigr)$ for the excess Bellman residual risk. Finally, under a concentrability condition, we translate this residual bound into an $L^2$ convergence rate of $\widetilde{\mathcal O}\bigl(n^{-\fracβ{d_x+d_a+2β}}\bigr)$ for the resulting deep estimator of $Q^*$, where $d_x$ and $d_a$ denote the dimensions of the state and action spaces, respectively, and $β$ denotes the Hölder smoothness index of $Q^*$. Importantly, our theoretical analysis does not rely on completeness assumptions commonly used in deep RL theory. Extensive numerical experiments demonstrate the effectiveness of the proposed method and its strong empirical performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑