arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RIFAR:面向持续机器人学习的可靠性与遗忘感知重放

RIFAR: Reliability and Forgetting-Aware Replay for Continual Robot Learning

Zirong Song, Zheng Lu, Haoran Liao, Wanqi Zhong, Yunhe Ni, Lijie Wang, Xiuying Chen

arXiv 2610.03079首次发表:更新:

发表机构

MBZUAI; Tsinghua University; Zhejiang University(穆罕默德·本·扎耶德人工智能大学; 清华大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RIFAR提出可靠性筛选与漂移感知重放方法,利用冻结逆动力学模型评估轨迹质量,在LIBERO基准上以极少历史数据超越现有生成式重放技术。

AI 中文摘要

真正的具身智能要求机器人将连续的真实世界经验转化为持久、可迁移的技能。这需要持续学习,在任务和环境演变时整合新能力而不侵蚀已有知识。经验重放可缓解遗忘,但随着任务累积,存储完整演示的成本变得高昂。世界-动作模型提供了一种生成式替代方案,通过联合预测动作和未来观测来重建过去的经验。然而,视觉上连贯的轨迹可能包含无法实现预测转换的动作,而新任务适应可能破坏先前学习的行为。因此,RIFAR将可靠性筛选与漂移感知重放选择相结合。它从紧凑的演示前缀重建轨迹,并使用冻结的逆动力学模型评估动作-视觉一致性。训练首先将当前演示与最高质量的筛选轨迹相结合。然后,RIFAR在相同历史输入上比较适应前后的动作预测,从同一筛选池中重新选择归一化漂移较大的轨迹以继续训练。在三个LIBERO套件和真实世界实验中,RIFAR超越了先前基于WAM的生成式重放的最先进水平。在LIBERO-Goal上,它达到了90.97的AUC,同时每个任务仅保留320个历史时间步,约为使用50个演示重放所保留步数的4.9%。

英文摘要

Genuine embodied agency requires robots to turn continuous real-world experience into lasting, transferable skills. This demands continual learning that integrates new capabilities without eroding prior knowledge as tasks and environments evolve. Experience replay mitigates forgetting, but storing complete demonstrations becomes costly as tasks accumulate. World-action models offer a generative alternative, reconstructing past experience through joint predictions of actions and future observations. However, visually coherent rollouts may contain actions that cannot realize the predicted transitions, while new-task adaptation can disrupt previously learned behavior. RIFAR therefore combines reliability screening with drift-aware replay selection. It reconstructs trajectories from compact demonstration prefixes and uses a frozen inverse-dynamics model to assess action-visual consistency. Training first combines current demonstrations with the highest-quality screened trajectories. RIFAR then compares action predictions before and after this adaptation on identical historical inputs, reselecting trajectories with larger normalized drift from the same screened pool for continued training. Across three LIBERO suites and real-world experiments, RIFAR surpasses the previous state of the art in WAM-based generative replay. On LIBERO-Goal, it achieves 90.97 AUC while retaining only 320 historical time steps per task, approximately 4.9% of the steps retained using 50-demonstration replay.

Comments15 pages, 6 figures, 9 tables, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑