arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

主动离线到在线强化学习

Active Offline-to-Online Reinforcement Learning

Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai

arXiv 2607.11720首次发表:更新:

发表机构

Virginia Commonwealth University; University of Virginia(弗吉尼亚共同体大学; 弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究离线到在线强化学习中有限交互预算下的主动策略选择问题,提出基于未来性能上置信界主动选策略微调的方法,经实验验证其优于现有基线,能更有效利用预算,推动离线强化学习在现实系统的实际部署。

AI 中文摘要

背景:离线强化学习能从大量先前收集的数据集中训练有效策略,并通过有限在线交互改进。这种离线到在线的强化学习范式在交互成本高或有潜在危险的非平稳领域很有前景。标准的离线到在线强化学习管道离线训练多个候选策略,用离策略或在线评估进行评估,然后部署并微调估计值最高的策略。然而,微调性能对算法和超参数选择高度敏感。目标:研究在离线到在线强化学习设置中有限交互预算下的主动策略选择以进行微调。方法:通过确定在将在线交互分配给策略评估(有助于识别高性能策略)和分配给微调(改善策略性能)之间的基本权衡来构建问题。然后提出一种方法,通过基于未来性能的上置信界主动选择策略进行微调来平衡这种权衡。这些界从拟合通过在线评估获得的观测值的局部线性性能预测中得出。结果:在各种实验中,该方法始终优于现有的离线到在线强化学习基线。结论:主动选择和微调策略比只采用单一策略或在所有策略间平均分配预算更有效地利用有限的在线交互预算。我们的框架还推动离线强化学习在在线交互成本高或有风险的现实世界系统中实际部署。

英文摘要

Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary domains where interaction is costly or potentially hazardous. Standard O2O-RL pipelines train multiple candidate policies offline, evaluate them using off-policy or online evaluation, and then deploy and fine-tune the policy with the highest estimated value. However, as in offline pretraining, fine-tuning performance is highly sensitive to the choice of algorithm and hyperparameters, making it risky to commit to a single policy. Objectives: We study active policy selection for fine-tuning under a limited interaction budget in O2O-RL settings. To our knowledge, this is the first work to address this problem. Methods: We formulate the problem by identifying a fundamental trade-off between allocating online interactions to policy evaluation, which helps identify high-performing policies, and allocating them to fine-tuning, which improves policy performance. We then propose an approach that balances this trade-off by actively selecting policies for fine-tuning based on upper-confidence bounds on their future performance. These bounds are derived from locally linear performance forecasts fitted to observations obtained through online evaluation. Results: Across a diverse range of experiments, the proposed approach consistently outperforms existing O2O-RL baselines. Conclusions: Actively selecting and fine-tuning policies uses limited online interaction budgets more effectively than either committing to a single policy or dividing the budget equally among all policies. Our framework also advances offline RL toward practical deployment in real-world systems where online interaction is costly or risky.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑