arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33186stat.MLcs.LG

基于混合贝尔曼残差与自适应评论家表示的离线策略评估

Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations

Amitakshar Biswas, Yuhan Li, Ruoqing Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种基于混合一步与两步贝尔曼残差及自适应评论家表示的离线策略评估方法,通过数据依赖的核表示和样本分割控制过拟合,在模拟与MetaWorld任务中验证了中间残差组合对值估计的改进效果。

中文摘要 AI 辅助

使用由不同行为策略生成的数据来评估目标策略,这仍然是强化学习中的一个基本挑战。尽管现有的大多数工作依赖于标准的一步贝尔曼残差,我们考虑了一步和两步残差的凸组合,并采用固定的混合权重。在理想情况下,这种混合贝尔曼公式可以在受限值函数类下的近似误差与由多步重要性加权引起的方差增加之间,提供一种自然的偏差-方差权衡。为了解决这种混合残差优化问题,我们采用了一种涉及评论家函数的极小极大公式。与依赖固定函数类的标准方法不同,我们使用预测的未来特征方向构建了一个数据依赖的评论家表示,这有效地诱导了一个适应底层转移动态的核。这使得评论家能够专注于与估计的贝尔曼误差最相关的方向。为了控制过拟合,我们使用样本分割来构建评论家,并在不同的数据子集上估计值函数。模拟研究和MetaWorld任务说明了混合参数的影响,并表明在具有挑战性的环境中,中间残差组合可以改善值估计。

英文摘要

Evaluating a target policy using data generated by a different behavior policy remains a fundamental challenge in reinforcement learning. While most existing work relies on the standard one-step Bellman residual, we consider a convex combination of one-step and two-step residuals with a fixed mixing weight. In ideal settings, this mixed Bellman formulation can provide a natural bias--variance trade-off between approximation error under a restricted value-function class and the increased variance arising from multi-step importance weighting. To solve this mixed residual optimization, we adopt a minimax formulation involving a critic function. Unlike standard approaches that rely on a fixed functional class, we construct a data-dependent critic representation using predicted future feature directions which effectively induces a kernel adapted to the underlying transition dynamics. This allows the critic to focus on directions that are most relevant for the estimated Bellman error. To control overfitting, we use sample splitting to construct the critic and estimate the value function on separate data subsets. Simulation studies and MetaWorld tasks illustrate the effect of the mixing parameter and show that intermediate residual combinations can improve value estimation in challenging settings.

发表机构

  • University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑