发表机构
University of California Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SAMBAR是一种持续学习算法,通过乘子法将学习视为约束优化,并选择性锚定关键参数,在无需旧任务演示的情况下防止VLA模型灾难性遗忘,平衡知识获取与保持。
AI 中文摘要
视觉-语言-动作(VLA)模型利用大规模预训练最终实现通用型操作。部署的VLA策略必须支持持续学习,以便随时间获取新任务。教授VLA新任务通常需要在其演示数据上进行微调。然而,对下游任务进行朴素微调会导致策略遗忘较早的任务,并降低通用能力。这种失败被称为灾难性遗忘。大多数持续学习方法通过重放早期任务的数据来应对这一问题。然而,旧任务的演示数据并不总是现成可得的。在本文中,我们引入了SAMBAR,一种持续学习算法,在VLA微调过程中防止灾难性遗忘,而无需访问任何先前学习任务的演示数据。我们提出将持续学习视为一个约束优化问题,并使用乘子法求解。在我们的方法中,乘子法驱动策略学习新任务,同时使模型参数不会偏离其先前值太远。与标准正则化惩罚不同,乘子法通过使用对偶变量,随着约束违反的累积而提高惩罚。我们还选择性地锚定对先前任务关键的参数以保留过去的知识,使其他参数自由用于新任务的获取。对偶变量与选择性锚定的结合因此平衡了知识获取与知识保持。我们在LIBERO仿真基准和硬件上评估了我们的方法SAMBAR。当在VLA上顺序微调时,我们比较的每个无重放基线都完全忘记了它学习的第一个任务,而SAMBAR保留了它学到的每一个任务。
英文摘要
Vision-Language-Action (VLA) models leverage large-scale pretraining to ultimately achieve generalist manipulation. Deployed VLA policies must support continual learning to acquire new tasks over time. Teaching a VLA a new task generally requires finetuning it on demonstrations of that task. However, naively finetuning on downstream tasks causes the policy to forget earlier tasks and degrades generalist capabilities. This failure is known as catastrophic forgetting. Most continual learning methods counter it by replaying data from earlier tasks. However, the old task demonstrations are not always readily available. In this paper, we introduce SAMBAR, a continual learning algorithm that prevents catastrophic forgetting during VLA finetuning without requiring access to the demonstrations of any previously learned task. We propose to cast continual learning as a constrained optimization problem and solve it with the method of multipliers. In our approach, the method of multipliers drives the policy to learn the new task without the model parameters drifting far away from their previous values. In contrast to a standard regularization penalty, the method of multipliers raises the penalty as the constraint violation accumulates by using a dual variable. We also selectively anchor the parameters critical to previous tasks to preserve past knowledge, leaving other parameters free for new task acquisition. The combination of dual variable and selective anchoring, therefore, balances knowledge acquisition with knowledge retention. We evaluate our method, SAMBAR, on the LIBERO simulation benchmark and on hardware. When sequentially finetuning on a VLA, every replay-free baseline we compare against completely forgets the first task it learned, whereas SAMBAR retains every task it has learned.