arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向类人物理智能:用于机器人操作的终身视觉-语言-动作学习

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation

Yao He, Gan Sun, Wenqi Liang, Fazeng Li, Yang Cong

arXiv 2607.14852首次发表:更新:

发表机构

South China University of Technology; University of Trento(华南理工大学; 特伦托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对机器人操作中可塑性-稳定性权衡难题,提出缓存高效的终身视觉-语言-动作学习框架LifelongVLA。通过双时间尺度LoRA门控模块及缓存高效随机重放策略,有效缓解权衡问题,实现技能扩展、行为保留,减少再训练依赖,优于现有基线。

AI 中文摘要

类似于人类依次学习新任务的自然能力,具有视觉-语言-动作(VLA)模型的机器人在开放世界环境中部署时应具备终身学习能力以学习新任务。然而,最近提出的大多数终身学习模型旨在有效学习当前任务(可塑性)或在先前任务上保持高精度(稳定性),而在机器人操作模型中可塑性-稳定性权衡在很大程度上仍未解决。为应对这一根本挑战,我们提出了一种用于机器人操作的缓存高效终身视觉-语言-动作学习框架(即LifelongVLA),它通过双时间尺度适应机制缓解可塑性-稳定性权衡,同时通过缓存高效重放策略实现低成本机器人部署。具体而言,我们提出了一个双时间尺度LoRA门控模块,将VLA适应分解为两条轻量级路径:一条用于可塑性的短期适配器和一条用于稳定巩固的长期适配器。这些路径通过任务感知门集成,实现对可塑性-稳定性权衡的显式控制。在技能重放阶段,提出了一种缓存高效随机重放策略,以在不进行全轨迹存储的情况下保留更平衡的保留信号。最后,实验表明LifelongVLA优于现有基线,展示了有效的技能扩展、对所学操作行为的稳健保留以及在xArm机器人上进行实际部署时对再训练的依赖减少。

英文摘要

Similar to the natural capabilities of humans to sequentially learn new tasks, robots with Vision-Language-Action (VLA) models should possess lifelong learning ability to learn a new task when deployed in open-world environments. However, most recently proposed lifelong learning models aim to effectively learn the current task (plasticity) or maintain high accuracy on previous tasks (stability), while the plasticity-stability trade-off remains largely unsolved in robotic manipulation models. To address this fundamental challenge, we propose a cache-efficient lifelong Vision-Language-Action learning framework for robotic manipulation (i.e., LifelongVLA), which alleviates the plasticity-stability trade-off with a dual-timescale adaptation mechanism while achieving low-cost robotic deployment with a cache-efficient replay strategy. More concretely, we propose a dual-timescale LoRA gating module to decompose VLA adaptation into two lightweight pathways: a short-term adapter for plasticity and a long-term adapter for stable consolidation. These pathways are integrated via a task-aware gate, enabling explicit control of the plasticity-stability trade-off. In the skill replay phase, a cache-efficient stochastic replay strategy is proposed to preserve more balanced retention signals without full-trajectory storage. Finally, experiments show that LifelongVLA outperforms existing baselines, demonstrating efficient skill expansion, robust retention of learned manipulation behaviors, and reduced reliance on retraining for real-world deployment on an xArm robot.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑