发表机构
Stanford University; Netflix; Cornell University(斯坦福大学; 网飞; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对离线强化学习中边缘化重要性加权的残留占用平衡违反问题,提出等渗贝尔曼校准方法,可减少此类违反并保证下游策略价值估计等任务的性能。
AI 中文摘要
边缘化重要性加权通过用目标策略的折扣占用比对离线状态-动作样本进行重加权来评估该策略,其特征是伴随贝尔曼方程。现有的极小极大、原始-对偶及拟合不动点估计器可能因函数类近似、正则化或优化不完整而留下残留的占用平衡违反情况。由于这些目标通常缺乏用于超参数调优、模型选择和早停的直接监督验证损失,因此难以诊断和减少此类违反情况。我们引入等渗贝尔曼校准,这是一种一维、与模型无关的后处理方法,可在保留任何初始占用比估计中的排序信息的同时减少上述违反情况。该方法通过在一维非递减变换类上应用拟合占用比评估(FORE)来校正估计的尺度和形状。我们将贝尔曼校准表征为条件不动点属性,等价于针对校准后比率的每个测试函数的占用平衡。更广泛地说,我们推导了校准-细化界,表明任何具有小校准误差的拟合比率的性能几乎与基于其拟合值的最佳后处理方法相当。对于等渗贝尔曼校准,我们建立了有限样本校准保证以及相对于初始估计的最佳单调变换的KL oracle不等式。因此,等渗贝尔曼校准实现了小的校准误差和KL风险,且处于最佳单调校正的统计误差范围内,同时为下游目标占用泛函(包括策略价值估计)提供保证。
英文摘要
Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equation. Existing minimax, primal-dual, and fitted fixed-point estimators can leave residual occupancy-balance violations because of function-class approximation, regularization, or incomplete optimization. These violations are difficult to diagnose and reduce because the objectives generally lack a direct supervised validation loss for hyperparameter tuning, model selection, and early stopping. We introduce isotonic Bellman calibration, a one-dimensional, model-agnostic post-processing method that reduces these violations while preserving the ranking information in any initial occupancy-ratio estimate. The method corrects the estimate's scale and shape by applying fitted occupancy-ratio evaluation (FORE) over a one-dimensional class of nondecreasing transformations. We characterize Bellman calibration as a conditional fixed-point property equivalent to occupancy-balance against every test function of the calibrated ratio. More generally, we derive a calibration-refinement bound showing that any fitted ratio with small calibration error performs nearly as well as the best post-processing based on its fitted values. For isotonic Bellman calibration, we establish finite-sample calibration guarantees and a KL oracle inequality relative to the best monotone transformation of the initial estimate. Consequently, isotonic Bellman calibration achieves small calibration error and KL risk within statistical error of the best monotone correction, with guarantees for downstream target-occupancy functionals, including policy-value estimation.
Comments43 pages, 1 figure, 4 tables