arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13730cs.LG

利用16万次训练运行的经验快速启动策略学习

JumpStart Your Policy Learning with Lessons from 160,000 Training Runs

Nabil Omi, Eric Bae, Chung Yik Edward Yeung, Siddhartha Sen, Ali Farhadi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过大规模实证(160,000+策略、114数据集)揭示离线策略学习中算法排名对超参数和基准的敏感性,并提出数据集条件推荐器及JumpStart资源套件以提升研究可靠性。

中文摘要 AI 辅助

离线策略学习的可靠进展依赖于细致的报告、调优良好的基线和跨多样条件的评估。先前工作表明,结果可能对报告选择、超参数调整和数据集属性敏感,但这些变异来源尚未在理解它们如何影响结论所需的规模上被系统地共同研究。为弥补这一空白,我们提出了一项大规模实证研究,涵盖离线强化学习和模仿学习,在114个数据集上训练了超过16万个策略。在此规模下,没有算法占据主导地位:最强方法之间的总体性能往往接近,但领先者在不同环境间差异显著。我们发现,适当的超参数调整经常重新洗牌感知的算法排名,且基准构成可能产生相互矛盾的结论。我们还研究了超参数敏感性和跨环境迁移,识别出一种推导强默认配置的简单策略。我们利用研究结果开发了一个数据集条件推荐器,为从业者提供特定任务的算法推荐。最后,我们发布了JumpStart:一个资源套件,包含所有训练过的策略、每个模型的得分和超参数、所有环境下的强基线、训练和评估代码,以及一个可扩展的网站,用于检索、分析和贡献结果。这些资源共同旨在使离线策略学习研究更加可靠,并支持超出本研究范围的未来工作。

英文摘要

Reliable progress in offline policy learning depends on careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work has shown that results can be sensitive to reporting choices, hyperparameter tuning, and dataset properties, but these sources of variability have not been systematically investigated together at the scale needed to understand how they shape conclusions. To address this gap, we present a large-scale empirical study of offline reinforcement and imitation learning, training over 160,000 policies across 114 datasets. At this scale, no algorithm dominates: aggregate performance among the strongest methods is often close, but the leaders differ substantially across environments. We find that proper hyperparameter tuning frequently reshuffles perceived algorithm rankings and that benchmark composition can produce conflicting conclusions. We also study hyperparameter sensitivity and transfer across environments, identifying a simple strategy for deriving strong default configurations. We use our findings to develop a dataset-conditioned recommender that provides task-specific algorithm recommendations for practitioners. Finally, we release JumpStart: a resource suite containing every trained policy, per-model scores and hyperparameters, strong baselines across all environments, training and evaluation code, and an extensible website for retrieving, analyzing, and contributing results. Together, these resources aim to make offline policy-learning research more reliable and enable future work beyond the scope of this study.

发表机构

  • University of Washington(华盛顿大学)
  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

↑