AI 中文总结
针对强化学习研究多基于模拟、向物理现实转化难的问题,提出开放蚂蚁平台。它是 Gymnasium 蚂蚁环境的物理变体及模拟,能让不同算法从物理机器人经验中快速学习行走策略,策略可迁移,还支持灵活实验,软硬件开源便于定制。
AI 中文摘要
强化学习研究在物理和模拟领域都取得了成功,但主要方法仍基于模拟。这使得算法和研究人员向物理现实转化研究存在不确定性。我们提出一个旨在简化这种转化的物理平台。本文介绍了开放蚂蚁,它是常用的 Gymnasium 蚂蚁环境的物理变体及模拟。我们证明,对于 SARSA(λ)和软演员评论家(SAC)这两种不同的强化学习算法,能直接从物理机器人经验中在约一小时内从头学习到胜任的行走策略。还展示了模拟中学习的策略可迁移到现实。此外,考察了该平台对灵活实验生态系统的支持情况,包括新用户取得首次成功的速度以及硬件问题出现时的修复和更新难易程度。硬件设计和软件在 GitHub 上开源以便定制。总之,我们倡导常使用模拟环境的强化学习研究人员使用开放蚂蚁,以便更轻松地在评估中纳入机器人实验。
英文摘要
Reinforcement learning (RL) research has demonstrated success in both physical and simulated domains; however, the predominant methodology remains rooted in simulations. The predominance of simulations makes translating research to physical reality uncertain for both algorithms and researchers. We propose a physical platform that is designed to simplify the transition. In this paper, we present the Open Ant: a physical variant of the commonly used Gymnasium Ant environment, along with a simulation. We demonstrate that competent walking policies can be learned from scratch in approximately one hour directly from the physical robot's experience for two substantially different RL algorithms: SARSA($λ$) and Soft Actor-Critic (SAC). Separately, we show policies that were learned in simulation transfer to reality. We also examine how well the platform supports a nimble experimental ecosystem. Specifically, we observe the speed with which new users from diverse backgrounds achieve their first success with the platform, and how easily the platform can be repaired and updated when hardware issues arise. Both the hardware design and software are available as open-source on GitHub for ease of customization. In summary, we advocate for the use of the Open Ant for RL researchers who frequently use simulated environments, so they can more easily include robot experiments in their evaluations.
CommentsPublished in the Reinforcement Learning Conference
Journal refReinforcement Learning Journal (RLJ), 2026