arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CIDER:面向具身强化学习的持续交互式蒸馏

CIDER: Continual Interactive Distillation for Embodied Reinforcement Learning

Houlin Li, Minghui Xu, Guo Xu, Xuan Du, Xiaohan Yan, Chun Wang, Yuxiang Yan, Shukai Yang, Yongcheng Liu, Wei Shan, Maoqing Yao

arXiv 2608.21899首次发表:更新:

发表机构

AgiBot; Shanghai Jiao Tong University(AgiBot; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出CIDER框架,通过冻结历史策略为教师策略并引入梯度路由,解决具身强化学习的灾难性遗忘问题,在6个现实世界操作任务上实现了新技能获取与旧技能保留的平衡。

AI 中文摘要

人机交互的现实世界强化学习能够在数十分钟内快速获取针对单个任务的有效机器人操作策略,然而目前仍不清楚如何将该范式扩展至持续学习场景,即单个策略需在获取新技能的同时不丢失先前学习的行为。现有现实世界持续学习方法未对先前行为施加显式约束,导致严重的灾难性遗忘。我们提出面向具身强化学习的持续交互式蒸馏框架(Continual Interactive Distillation for Embodied Reinforcement Learning,简称CIDER),该框架在学习每个新任务前将累积的历史策略冻结为教师策略,并在任务学习过程中穿插基于蒸馏的保留操作。我们进一步引入梯度路由机制,将用于获取新任务的梯度与用于保留先前行为的梯度分离开来。我们在6个现实世界家庭及工业操作任务上采用单个共享执行器对我们的方法进行评估。交互式蒸馏在6任务现实机器人序列中对先前学习任务保持了较高的测量成功率,且每个新任务的获取耗时为10至20分钟,而所有基准方法均至少遗忘了一个先前任务。额外的 ablation 研究揭示了决定现实世界持续强化学习中稳定性与可塑性之间权衡的关键设计选择。

英文摘要

Human-in-the-loop real-world reinforcement learning enables rapid acquisition of effective robotic manipulation policies for individual tasks, often within tens of minutes. Yet it remains unclear how to extend this paradigm to continual learning, where a single policy must acquire new skills without losing previously learned behaviors. Existing real-world continual learning methods do not explicitly constrain prior behaviors, leading to severe catastrophic forgetting. We introduce Continual Interactive Distillation for Embodied Reinforcement Learning (CIDER), a continual reinforcement learning framework that freezes the accumulated historical policy as a teacher before learning each new task and interleaves task learning with distillation-based retention. We further introduce gradient routing to separate the gradients used for acquiring new tasks from those used for preserving prior behaviors. We evaluate our method with a single shared actor on six real-world household and industrial manipulation tasks. Interactive Distillation maintains high measured success on previously learned tasks across our six-task real-robot sequence while acquiring each new task in 10 to 20 minutes, whereas every baseline forgets at least one previous task. Additional ablations reveal the key design choices that govern the tradeoff between stability and plasticity in real-world continual reinforcement learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑