发表机构
Mila - Quebec AI Institute; Université de Montréal; McGill University; Google DeepMind(Mila - 魁北克人工智能研究所; 蒙特利尔大学; 麦吉尔大学; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出离线强化学习应学习自适应策略先验而非静态策略,通过记忆、探索和自我纠错在后续交互中保持改进能力,并讨论贝叶斯方法作为构建方向。
AI 中文摘要
离线强化学习(RL)传统上侧重于在保守目标下学习可直接部署的策略,其中对离线数据集之外的不确定性采取悲观处理以确保鲁棒性。我们认为,当离线训练的策略随后通过在线交互进行更新时(这在现代智能系统中通过测试时适应和在线微调日益普遍),这种表述变得不完整。本立场论文认为,在此类设置中,离线强化学习的目标应超越即时部署,优先学习自适应策略先验:即通过记忆、探索和自我纠错在后续交互中保持改进能力的策略。我们将这一视角形式化为自适应离线强化学习(AORL),将其与离线到在线强化学习区分开来,并解释为何在分布偏移、有限数据集覆盖和变化的测试时条件下适应性变得重要。我们进一步讨论贝叶斯离线强化学习作为构建自适应策略先验的一个原则性方向,通过保留对可能环境的认知不确定性。最后,我们概述了将离线强化学习视为对未来经验的准备而非静态部署问题的联系、开放挑战和研究方向。
英文摘要
Offline reinforcement learning (RL) has traditionally focused on learning policies for direct deployment under conservative objectives, where uncertainty outside the offline dataset is treated pessimistically to ensure robustness. We argue that this formulation becomes incomplete when an offline-trained policy is subsequently updated through online interaction, as increasingly occurs in modern intelligent systems through test-time adaptation and online fine-tuning. This position paper argues that, in such settings, the objective of offline RL should extend beyond immediate deployment and instead prioritize learning adaptive policy priors: policies that preserve the capacity to improve during subsequent interaction through memory, exploration, and self-correction. We formalize this perspective as adaptive offline reinforcement learning (AORL), distinguish it from offline-to-online RL, and explain why adaptability becomes important under distributional shift, limited dataset coverage, and changing test-time conditions. We further discuss Bayesian offline RL as one principled direction for constructing adaptive policy priors by preserving epistemic uncertainty over plausible environments. Finally, we outline connections, open challenges, and research directions for treating offline RL as preparation for future experience rather than as a static deployment problem.
CommentsAccepted to NeurIPS Position Paper Track, 2026