发表机构
Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对离线强化学习面临的挑战,提出不确定性估计下的保守查询与自适应正则化框架,通过莫尔斯网络估计不确定性,采用保守查询策略与自适应正则化方案,与CQL集成并在D4RL基准测试,取得优异或有竞争力的性能。
AI 中文摘要
离线强化学习旨在从静态数据集中学习有效策略,但其性能受数据集覆盖限制。动作偏好查询可利用专家反馈且无需额外环境交互来改进离线训练策略。现有方法面临选择信息丰富的偏好查询及有效利用收集到的反馈这两个关键挑战。本文提出不确定性估计下的保守查询与自适应正则化框架,它通过莫尔斯网络估计策略动作关于离线数据集的不确定性,引入保守查询策略和不确定性感知自适应正则化方案。该框架与CQL集成并在D4RL基准上广泛评估,实验结果显示在众多任务中性能优异或具有竞争力。
英文摘要
Offline reinforcement learning (RL) aims to learn an effective policy from a static dataset, but its performance is fundamentally limited by dataset coverage. Action preference queries leverage expert feedback without additional environment interaction, enabling policy improvement during offline training. However, existing methods still face two key challenges: selecting informative preference queries and effectively exploiting the collected feedback. Current approaches typically rely only on the distance between policy actions and dataset actions for query selection, while enforcing fixed constraints that keep the policy close to queried preferences. Such strategies often lead to unstable policy updates and integrate poorly with value regularization. To address these limitations, we propose Conservative Query and Adaptive Regularization under Uncertainty Estimation, a lightweight framework that jointly improves preference querying and preference exploitation. Specifically, we employ a Morse network to estimate the uncertainty of policy actions with respect to the offline dataset. Based on this uncertainty, we introduce a conservative query strategy that selectively queries actions near the dataset to preserve Bellman-update stability, together with an uncertainty-aware adaptive regularization scheme that dynamically adjusts data-level constraints during policy optimization. We integrate our framework with CQL and evaluate it extensively on the D4RL benchmark. Experimental results demonstrate superior or competitive performance across a wide range of tasks.
CommentsAccepted by ECAI2025