流式深度强化学习用于机器人自适应持续学习的分析
An Analysis of Streaming Deep Reinforcement Learning for Adaptive Continual Learning in Robotics
- Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文首次分析流式深度强化学习在机器人持续学习中的应用,证明其能使机器人适应未预见变化,在四足运动中成功率提升高达90%,并在操作任务中部分有效。
AI中文摘要:
在机器人的生命周期中,可能会遇到其原始训练中未预料到的新场景,导致性能下降。缓解这一问题的一种常见方法是进一步扩充离线训练数据集,以期产生对这些变化具有鲁棒性的策略。相比之下,生物学习是通过经验流逐时刻发生的,这与深度学习主要基于批次和离线的性质不同。尽管近期研究表明基于流的深度强化学习(其中更新仅使用最新经验)的可行性,但尚未有研究证明其作为持续学习框架能够适应机器人策略以应对未见变化。在本文中,我们首次对用于机器人自适应持续学习的流式深度强化学习进行了分析。具体而言,我们表明,在初始预训练阶段之后,流式深度RL可以使机器人成功适应自身、环境或目标的未预见变化。我们在四足运动中的主要实验表明,采用特定优化器和塑性损失缓解技术的深度神经网络机器人策略能够利用预训练中的领域任务知识,通过流学习快速在线适应各种变化,优于基于批次的在线策略方法,并将任务成功率比预训练策略提高高达90%。此外,我们对机器人操作任务进行了额外评估,以确定我们之前的观察是否适用于不同的机器人形态和场景。我们的结果表明,在四足运动中观察到的成功可以在操作任务中部分实现,但存在稳定性和性能限制。最后,我们讨论了工作的局限性及其对未来持续机器人学习的影响。
英文摘要:
Over the course of a lifetime, robots may encounter novel scenarios unaccounted for in its original training that result in performance degradation. One common approach to mitigating this issue is to further grow the offline training dataset in hopes of producing a policy robust to these changes. In contrast, biological learning occurs moment-to-moment via a stream of experience, unlike the predominantly batch-based and offline nature of deep learning. Although recent works show the feasibility of stream-based deep reinforcement learning, where updates use only the latest experience, none have shown it to be a viable continual learning framework for adapting robotic policies to unseen changes. In this paper, we present the first analysis of streaming deep reinforcement learning for adaptive continual learning in robotics. In particular, we show that, following an initial pretraining phase, streaming deep RL can enable a robot to successfully adapt to unforeseen changes to itself, its environment, or goals. Our primary experiments within quadruped locomotion demonstrate that a deep neural network robotic policy with certain optimizers and plasticity loss mitigation techniques can successfully leverage domain task knowledge from its pretraining to quickly adapt online to diverse changes via stream learning, outperforming batch-based on-policy methods and improving task success rates by up to 90% over the pretrained policy. Furthermore, we perform additional evaluations on robotic manipulation tasks to determine if our previous observations extend to different robotic morphologies and scenarios. Our results show that the successes observed in quadruped locomotion can be partially realized in manipulation with stability and performance limitations. We conclude with a discussion on the limitations of our work and its implications for the future of continual robot learning.