发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对离线强化学习分布偏移下的离策略评估问题,提出拟合占用比评估方法FORE,通过伴随贝尔曼递归估计占用比,仅需占用比可实现性,无需贝尔曼完备性等条件,保障离线策略评估收敛性。
AI 中文摘要
占用比可校正离线强化学习中的分布偏移,是离策略评估的核心要素。现有原始-对偶与极小极大方法通常通过在评论函数类上施加占用平衡矩约束来估计占用比。本文提出拟合占用比评估(FORE)方法,这是一种拟合不动点方法,可通过伴随贝尔曼递归刻画折扣占用比。在每一轮迭代中,FORE针对单步转移数据求解单层密度比目标,从而将伴随贝尔曼像以KL散度为度量投影到对数比函数类上。与通常要求值函数可实现性搭配贝尔曼完备性或投影算子稳定性的拟合Q评估分析不同,本文的核心近似条件仅为折扣占用比自身的可实现性。在该条件下,由于伴随贝尔曼算子是KL压缩算子,总体KL投影递归会在相对熵意义下朝着真实占用比收缩。针对经验递归过程,本文推导了有限样本遗憾界,可保证算法在对数比近似误差与由比假设类复杂度决定的统计误差范围内实现KL收敛。拟合得到的占用比可通过奖励重加权、占用加权拟合Q评估,以及结合拟合占用比与拟合Q函数的双稳健估计实现直接值估计。上述结果共同证明,折扣占用比可实现性是无需任何完备性假设即可开展离线策略评估的充分条件。
英文摘要
Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class. We propose fitted occupancy-ratio evaluation (FORE), a fitted fixed-point method that characterizes the discounted occupancy ratio through an adjoint Bellman recursion. At each iteration, FORE solves a single-level density-ratio objective on one-step-transition data, thereby projecting the adjoint Bellman image onto a log-ratio class in Kullback-Leibler (KL) divergence. Unlike analyses of fitted Q-evaluation, which typically require value-function realizability together with Bellman completeness or projected-operator stability, our central approximation condition is just realizability of the discounted occupancy ratio itself. Under this condition, the population KL-projected recursion contracts in relative entropy toward the true ratio by virtue of the adjoint Bellman operator being a KL-contraction. For the empirical recursion, we establish finite-sample regret bounds that yield convergence in KL up to approximation error and a statistical error governed by the complexity of the ratio hypothesis class. When full coverage fails, we introduce coverage-stopped FORE, which targets the discounted occupancy accumulated before the first uncovered state-action pair and yields a conservative lower bound on target-policy value for nonnegative rewards. The fitted ratio supports direct value estimation by reward reweighting, occupancy-weighted fitted Q-evaluation, and doubly robust estimation that combines the fitted ratio with a fitted Q-function. Together, these results identify discounted occupancy-ratio realizability as a sufficient condition for offline policy evaluation without any completeness assumptions.