AI 中文总结
针对离线强化学习中分布外动作估计误差被放大的问题,提出CSDG算法,通过分离样本内参考与局部修正项实现泛化控制,经实验在Gym-MuJoCo等任务上表现强劲。
AI 中文摘要
离线强化学习(offline RL)可从附近的分布外(OOD)动作中获益,但这些动作处的估计误差可能会引导自举(bootstrapping)被放大。现有正则化和局部泛化方法通常通过独立机制控制可允许的OOD区域或泛化目标的影响。我们提出凸包邻域平滑对偶泛化(Convex-Hull-Neighborhood Smooth Dual Generalization,CSDG),该方法将贝尔曼备份表示为样本内价值目标加上CHN局部修正。此公式明确了泛化贡献,并将其与样本内参考路径分离。修正项通过在不同扰动半径下采样的样本内导向和OOD导向候选进行平滑得到,混合系数λ缩放其对每次备份的贡献,同时递归折扣仍为γ。在有界性和固定扰动核的条件下,我们推导了精确的单步修正恒等式、时变迭代界以及仅依赖于不动点处分支差异的不动点界。我们进一步刻画了由理想化算子诱导的隐式策略,并给出了条件非退化准则。该实用算法通过非对称有界噪声和期望回归近似这些量,无需精确的支持分类或额外的悲观OOD惩罚。在Gym-MuJoCo和AntMaze上的实验显示出强劲的综合性能和稳定的价值估计,代码可在this https URL获取。
英文摘要
Offline reinforcement learning (offline RL) can benefit from nearby out-of-distribution (OOD) actions, but estimation errors at these actions may be amplified by bootstrapping. Existing regularization and local-generalization methods control either the admissible OOD region or the influence of generalized targets, often through separate mechanisms. We propose Convex Hull Neighborhood Smooth Dual Generalization (CSDG), which expresses the Bellman backup as an in-sample value target plus a CHN-local correction. This formulation makes the generalized contribution explicit and separates it from the in-sample reference path. The correction is obtained by smoothing in-sample-oriented and OOD-oriented candidates sampled at different perturbation radii. A mixture coefficient lambda scales its contribution to each backup, while the recursive discount remains gamma. Under boundedness and fixed perturbation kernels, we derive an exact one-step correction identity, a time-varying iterate bound, and a fixed-point bound that depends only on the branch discrepancy at the fixed point. We further characterize the implicit policies induced by the idealized operators and give a conditional non-degradation criterion. The practical algorithm approximates these quantities using asymmetric bounded noise and expectile regression, without exact support classification or an additional pessimistic OOD penalty. Experiments on Gym-MuJoCo and AntMaze show strong aggregate performance and stable value estimation. Code is available at: https://github.com/YOUNG-fnxm/CSDG