arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考实际迁移约束下强化学习算法的适用性

Rethinking the Suitability of Reinforcement Learning Algorithms Under Practical Transfer Constraints

Hany Hamed, Abhishek Naik, Colin Bellinger, A. Rupam Mahmood

arXiv 2607.17326首次发表:更新:

AI 中文总结

研究实际迁移约束下强化学习算法适用性,从实际效率和动态不匹配鲁棒性两维度评估。发现PPO虽样本效率低但训练快,域随机化对不同算法影响相似,强调评估算法不能仅看样本效率,还需考虑实际因素。

AI 中文摘要

面向迁移的强化学习需要从超越标准样本效率的维度评估算法。我们关注两个维度:实际效率,即在时钟时间而非基于交互的预算下,算法适用性结论是否改变;动态不匹配下的鲁棒性,即不同学习范式如何应对由域随机化引起的训练分布变化。我们为强化学习从业者提供两点见解。首先,在面向迁移的设置中,比较不同算法的样本效率往往不够。训练一个合适策略所需的时钟时间对从业者很重要,我们发现样本效率低的PPO算法能比样本效率相对高的算法(如SAC和TD-MPC2)更快产生高性能策略。其次,域随机化可帮助不同算法学习鲁棒策略。我们发现域随机化对PPO、SAC和TD-MPC2的影响相似。这两点见解突出了不仅按样本效率,还按训练时间等实际因素评估强化学习算法的重要性。

英文摘要

Transfer-oriented reinforcement learning requires evaluating algorithms along dimensions that go beyond standard sample efficiency. We focus on two dimensions: practical efficiency, which asks whether conclusions about algorithm suitability change under wall-clock rather than interaction-based budgets, and robustness under dynamics mismatch, which asks how different learning paradigms respond to variability in the training distribution induced by domain randomization. We provide two insights to reinforcement-learning practitioners. First, comparing the sample efficiency of different algorithms is often an insufficient criterion in transfer-oriented settings. The wall-clock time required to train a decent policy is an important consideration for practitioners, and we find that the sample-inefficient PPO algorithm can produce a performant policy faster than relatively more sample-efficient algorithms such as SAC and TD-MPC2, validating the common understanding of massively parallel training paradigms. Second, domain randomization can help different kinds of algorithms learn robust policies. In particular, although PPO, SAC, and TD-MPC2 represent different RL paradigms - on-policy, off-policy, and model-based learning and planning, respectively - we find that domain randomization affects all three algorithms in a similar way. To the best of our knowledge, this is the first controlled comparison of the effect of domain-randomization coverage on PPO, SAC, and TD-MPC2 under the same transfer protocol. Taken together, these two insights highlight the importance of evaluating RL algorithms not only by sample efficiency, but also by practical considerations such as training time and the algorithms' ability to produce usable policies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑