arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34593cs.LGstat.ML

深度加权贝尔曼残差最小化用于 $Q^*$ 估计

Deep Weighted Bellman Residual Minimization for $Q^*$ Estimation

发表机构武汉大学
查看机构详情
  • Wuhan University(武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

Lican Kang, Jerry Zhijian Yang, Cheng Yuan, Chen Zhong

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出深度加权贝尔曼残差最小化框架,通过密度比加权整合专家演示与行为数据,解决离线强化学习中的分布偏移和Q值高估问题,并显著提升数值性能与策略泛化。

中文摘要 AI 辅助

离线策略评估是离线强化学习的基础组成部分,旨在利用预先收集的数据集评估和优化策略性能。然而,此类数据集常常面临显著挑战,包括分布偏移、$Q$ 值高估以及样本利用效率低下。为解决这些问题,本文引入了一种加权贝尔曼残差最小化框架,该框架通过有效整合专家演示与行为数据,纳入密度比加权。所提出的加权方案偏离了深度强化学习理论分析中通常施加的常规完备性假设。我们建立了密度比估计的尖锐收敛速率,并推导了所得深度 $Q^*$ 估计器超额风险的收敛速率。广泛的实证评估表明,与现有方法相比,我们的方法在数值性能和策略泛化方面取得了显著改进,为专家演示的合理利用提供了具体指导。

英文摘要

Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, $Q$-value overestimation, and low sample utilization efficiency. To address these issues, this paper introduces a weighted Bellman residual minimization framework that incorporates density ratio weighting by effectively integrating expert demonstrations with behavioral data. The proposed weighting scheme departs from the conventional completeness assumption commonly imposed in the theoretical analysis of deep reinforcement learning. We establish a sharp convergence rate for density ratio estimation and derive the convergence rate for the excess risk of resulting deep $Q^*$ estimator. Extensive empirical evaluations demonstrate that, compared to existing methods, our method achieves significant improvements in numerical performance and policy generalization, providing specific guidance for the rational utilization of expert demonstrations.

↑