arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20389stat.MLcs.LG

基于模型的Bootstrap方法用于表格强化学习中的离线策略评估

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

Weiwei Wang, Yuqiang Li, Xianyi Wu, Bingyi Jing

首次发表
浏览论文内容

中文总结 AI 辅助

针对离线策略评估中的不确定性量化问题,提出基于模型的Bootstrap框架,通过从估计的MDP重生成轨迹,适应多种数据格式,实现分布一致性和有效置信区间,提升有限样本效率。

中文摘要 AI 辅助

离线策略评估(OPE)在高风险强化学习应用中至关重要,在这些应用中,新策略必须在部署前得到可靠评估。在此类场景中,仅有点估计是不够的;原则性的不确定性量化,如置信区间和方差估计,对于安全和风险感知的决策至关重要。统一这些任务的一个全面方法是估计评估误差的采样分布。然而,现有方法常常在鲁棒性、可扩展性或有限样本有效性方面存在不足。在本文中,我们提出了一种基于模型的Bootstrap框架,用于有限时域、非时齐马尔可夫决策过程(MDPs)中OPE的不确定性量化。与依赖重采样完整情节的经典Bootstrap方法不同,所提出的方法从估计的MDP中重新生成轨迹,因此能够适应更广泛的离线数据格式,包括完整轨迹、转移级观测和轨迹片段。这种灵活性进一步提高了有限样本的统计效率。我们建立了Bootstrap分布一致性、渐近有效的置信区间以及目标策略值的一致方差估计。大量模拟表明,所提出的方法能够准确捕捉OPE估计器的采样分布,在大多数设置中产生更紧的置信区间和更准确的方差估计。

英文摘要

Offline policy evaluation (OPE) is crucial in high-stakes reinforcement learning applications, where new policies must be assessed reliably before deployment. In such settings, point estimates alone are insufficient; principled uncertainty quantification, such as confidence intervals and variance estimates, is essential for safe and risk-aware decision-making. A comprehensive way to unify these tasks is to estimate the sampling distribution of the evaluation error. Existing approaches, however, often suffer from limited robustness, scalability, or finite-sample validity. In this paper, we propose a model-based bootstrap framework for uncertainty quantification of OPE in finite-horizon, time-inhomogeneous Markov decision processes (MDPs). Unlike classical bootstrap methods that rely on resampling complete episodes, the proposed method regenerates trajectories from an estimated MDP and can therefore accommodate a much broader range of offline data formats, including complete trajectories, transition-level observations, and trajectory fragments. This flexibility further improves finite-sample statistical efficiency. We establish bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation for the target policy value. Extensive simulations show that the proposed method accurately captures the sampling distribution of the OPE estimator, yielding tighter confidence intervals and more accurate variance estimates in most settings.

发表机构

  • Southern University of Science and Technology(南方科技大学)
  • East China Normal University(华东师范大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Shenzhen Loop Area Institute(深圳河套学院)

机构由 AI 辅助整理,请以论文原文为准。

↑