arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.00338cs.LGcs.AI

高吞吐环境下的可扩展选项学习

Scalable Option Learning in High-Throughput Environments

  • Meta Superintelligence Labs(Meta超智能实验室)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Mikael Henaff, Scott Fujimoto, Michael Matthews, Michael Rabbat

更新

AI总结:

本文提出SOL算法,通过高吞吐量的层级强化学习方法,在NetHack等复杂游戏中实现35倍的吞吐量提升,验证了其在高吞吐环境中的扩展性和通用性。

AI中文摘要:

层级强化学习(RL)有潜力在长时间尺度上实现有效的决策。现有方法虽然有前景,但尚未实现大规模训练的好处。本文识别并解决了将在线层级RL扩展到高吞吐环境中的几个关键挑战。我们提出可扩展选项学习(SOL),一种高度可扩展的层级RL算法,其吞吐量比现有层级方法高约35倍。为了展示SOL的性能和可扩展性,我们在复杂的NetHack游戏中使用300亿帧的经验训练层级代理,显著超越了扁平代理,并展示了积极的扩展趋势。我们还在MiniHack和Mujoco环境中验证了SOL,展示了其通用性。我们的代码已开源:github.com/facebookresearch/sol.

英文摘要:

Hierarchical reinforcement learning (RL) has the potential to enable effective decision-making over long timescales. Existing approaches, while promising, have yet to realize the benefits of large-scale training. In this work, we identify and solve several key challenges in scaling online hierarchical RL to high-throughput environments. We propose Scalable Option Learning (SOL), a highly scalable hierarchical RL algorithm which achieves a ~35x higher throughput compared to existing hierarchical methods. To demonstrate SOL's performance and scalability, we train hierarchical agents using 30 billion frames of experience on the complex game of NetHack, significantly surpassing flat agents and demonstrating positive scaling trends. We also validate SOL on MiniHack and Mujoco environments, showcasing its general applicability. Our code is open sourced at: github.com/facebookresearch/sol.

↑