arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于RoboRacer平台自主赛车中泛化的持续强化学习

Continual-RL for Generalization in Autonomous Racing on the RoboRacer Platform

Joel Siegert, Edoardo Ghignone, Michele Magno

arXiv 2607.24320首次发表:更新:

发表机构

Center for Project-Based Learning, D-ITET, ETH Zurich(基于项目的学习中心,分布式信息技术与电气工程学院,苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现代机器人适应变化环境的挑战,提出基于持续反向传播的持续强化学习框架,仅用现实世界数据训练通用策略并微调超越经典控制器,还提出离线强化学习比较方法并进行模拟分析。

AI 中文摘要

现代机器人技术的一个关键挑战是适应不断变化的环境,当模拟无法涵盖所有可能的现实世界配置时,这一挑战会加剧,因此物理世界中的强化学习变得必要。持续强化学习提供了应对这一挑战的工具,但框架和方法仍未得到充分探索。自主赛车,特别是RoboRacer竞赛为这类方法提供了试验场。本文提出了一种基于持续反向传播的持续强化学习框架,仅使用现实世界数据就能在一组赛道上训练通用策略,并在15分钟内对其进行微调以超越经典控制器。此外,还提出了一种基于离线强化学习的比较方法,并对这些方法的可塑性特性进行了模拟分析。

英文摘要

A key challenge in modern robotics is to adapt to changing environments, a challenge that is exacerbated when simulations cannot encompass every possible real-world configuration, and therefore Reinforcement Learning (RL) in the physical world becomes necessary. Continual Reinforcement Learning provides the tools to address this challenge; however, both the frameworks and the methods remain underexplored. Autonomous Racing and in particular the RoboRacer competition provide a testing ground for such methods, as learning to drive on a new track-floor combination with the least amount of new experience naturally frames a continual learning problem. This work tries to address this gap by proposing a continual RL framework based on Continual Backpropagation that is able, with only real-world data, to train a generalistic policy on a set of tracks and then fine- tune it within 15 minutes to outperform classical controllers. Furthermore, a comparison method based on offline RL is proposed, and a simulation analysis of the plasticity properties of the methods is conducted.

Comments8 pages, conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑