arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GzDRL:基于Gazebo的可复现且可扩展的深度强化学习

GzDRL: Reproducible and Scalable Deep Reinforcement Learning with Gazebo

Amal Dev Haridevan, Junjie Kang, Jinjun Shan

arXiv 2609.13243首次发表:更新:

AI 中文总结

GzDRL提出无中间件的单进程RL框架,实现Gazebo中确定性、高吞吐、可复现的训练,并验证了四旋翼的仿真到现实迁移。

AI 中文摘要

我们提出了GzDRL,一个用于Gazebo的新型单进程强化学习(RL)框架,它克服了可扩展、可复现机器人实验中长期存在的瓶颈。与传统的基于中间件的RL-Gazebo集成因非确定性和不可复现性而受影响不同,GzDRL引入了一种系统的、无中间件的环境步进机制,直接同步智能体动作和物理更新。这种设计实现了确定性的、高吞吐量的数据收集、高效的向量化以及可复现的RL训练和评估。综合基准测试表明,GzDRL在评估的框架中实现了最高的工作站吞吐量,同时在笔记本电脑硬件上与GPU加速模拟器保持竞争力,并保持了精确的智能体-环境同步、多智能体可扩展性和实验级别的可复现性。我们通过将学习到的策略直接部署到物理四旋翼飞行器上,无需微调,进一步验证了仿真到现实的迁移。我们的结果确立了GzDRL作为推进机器人和自动化领域RL的可访问且可复现的平台。

英文摘要

We present GzDRL, a novel single-process reinforcement learning (RL) framework for Gazebo that overcomes longstanding bottlenecks in scalable, reproducible robotics experimentation. Unlike conventional middleware-based RL-Gazebo integrations that suffer from nondeterminism and irreproducibility, GzDRL introduces a systematic, middleware-free environment-stepping mechanism that directly synchronizes agent actions and physics updates. This design enables deterministic, high-throughput data collection, efficient vectorization, and reproducible RL training and evaluation. Comprehensive benchmarks demonstrate that GzDRL achieves the highest workstation throughput among the evaluated frameworks while remaining competitive with GPU-accelerated simulators on laptop hardware, and maintains precise agent-environment synchronization, multi-agent scalability, and experiment-level reproducibility. We further validate sim-to-real transfer by deploying learned policies directly onto a physical quadrotor, without fine-tuning. Our results establish GzDRL as an accessible and reproducible platform for advancing RL in robotics and automation.

CommentsThis work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑