内联记忆与可复用技能:面向视觉-语言-动作模型的记忆中心框架
Inline Memory Meets Reusable Skills: Memory-centric Framework for Vision-Language-Action Model
- Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
- Pengcheng Laboratory(鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对VLA模型适配低效与灾难性遗忘问题,提出记忆中心框架Optimus-R,通过内联记忆接口、查询-技能记忆库及桥接-适配机制实现数据高效技能学习,实验验证其有效性。
AI中文摘要:
视觉-语言-动作(VLA)模型在通用机器人操作任务中展现出巨大潜力,但将其适配到新任务和新领域仍然效率低下:现有方法往往依赖参数微调,导致成本高昂,并存在灾难性遗忘先前学习任务的风险。为解决这一问题,我们提出Optimus-R,一种以记忆为中心的VLA框架,将机器人适配形式化为显式的查询-技能记忆微调。Optimus-R引入:(i)用于技能提取的内联记忆接口。它在VLA前缀流中插入可学习的记忆令牌,使骨干网络能够在原生动作条件路径中推导出控制感知的查询和技能表示。(ii)用于技能学习的查询-技能记忆库。它将技能外部化为查询原型以决定检索内容,以及技能值以指定执行方式,支持技能复用和扩展,同时仅需有限的参数更新。(iii)用于技能更新的轻量级桥接-适配机制。它通过轻量级适配器和残差记忆更新,将目标域查询和技能与现有记忆空间对齐。在域内适配、跨域适配和终身学习实验中的结果表明,Optimus-R能够实现数据高效的技能学习,同时缓解灾难性遗忘。
英文摘要:
Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains remains inefficient: existing methods often rely on parameter tuning, incurring substantial costs and risking catastrophic forgetting of previously learned tasks. To address this, we propose \textbf{Optimus-R}, a memory-centric VLA framework that formulates robotic adaptation as explicit query-skill memory tuning. Optimus-R introduces: (i) An \textbf{Inline Memory Interface for skill extraction}. It inserts learnable memory tokens into the VLA prefix stream, allowing the backbone to derive control-aware query and skill representations within the native action-conditioning pathway. (ii) A \textbf{Query-Skill Memory Bank for skill learning}. It externalizes skills into query prototypes for deciding \emph{what} to retrieve and skill values for specifying \emph{how} to act, supporting skill reuse and expansion with limited parameter updates. (iii) A lightweight \textbf{Bridge-and-Adapt mechanism for skill updating}. It aligns target-domain queries and skills with the existing memory space through a lightweight adapter and residual memory updates. Experiments on in-domain adaptation, cross-domain adaptation, and lifelong learning show that Optimus-R enables data-efficient skill learning while mitigating catastrophic forgetting.