OrthoSkillVLA:通过梯度感知技能子空间自适应实现持续技能学习
OrthoSkillVLA: Continual Skill Learning via Gradient-Informed Skill Subspace Adaptation
浏览论文内容
中文总结 AI 辅助
OrthoSkillVLA是一种参数高效的持续技能学习框架,通过对VLM和ActionHead施加独立子空间约束、引入特征感知MoE解码器,实现预训练VLA模型的技能学习,可缓解灾难性遗忘。
中文摘要 AI 辅助
预训练视觉-语言-动作(Vision-Language-Action,VLA)模型为机器人学习提供了坚实基础,但将其依次适配到不同技能时,会干扰先前技能所用的表征和速度映射,导致灾难性遗忘。基于架构的方法通过隔离技能来提升保留能力,但会增加推理开销;近期的子空间约束方法限制正交子空间内的参数更新以最小化干扰,但对整个模型施加统一约束。我们分析了VLA内部组件的不同作用,识别出两个VLA特有的挑战:第一,视觉语言模型(Vision-Language Model,VLM)维持广泛的语义表征,易出现容量耗尽,而动作头(ActionHead)将语义细化为局部速度模式,对扰动高度敏感;第二,最终速度解码器作为读出层,冻结它会形成输出阶段的表达瓶颈,更新它则有覆盖先前速度映射的风险。为此,我们提出OrthoSkillVLA,一种用于预训练VLA模型持续技能学习的参数高效框架,无需演示回放。鉴于表征异质性,我们对VLM和ActionHead施加独立的子空间约束,在保留可复用语义容量的同时保护局部速度模式;对于输出层,我们引入轻量的特征感知MoE(混合专家模型)解码器,每个技能分配一个紧凑专家,无训练的路由器根据特征空间相似度选择专家。大量模拟与真实世界评估及 ablation 实验表明,OrthoSkillVLA在获取新技能的同时能更好地保留先前技能。
英文摘要
Pretrained Vision-Language-Action models provide a strong foundation for robot learning, but sequentially adapting them to diverse skills can perturb the representations and velocity mappings used by previous skills, leading to catastrophic forgetting. Architecture-based approaches improve retention by isolating skills but lead to increased inference footprint. Recent subspace-constrained methods restrict parameter updates in an orthogonal subspace to minimize interference but impose a unified constraint on the entire model. We analyze the distinct roles of internal VLA components and identify two VLA-specific challenges. First, the VLM maintains broad semantic representations, making it vulnerable to capacity exhaustion, whereas the ActionHead refines semantics into localized velocity patterns that are highly sensitive to perturbations. Second, the final velocity decoder serves as a readout layer. Freezing it forms an output-stage expressivity bottleneck, while updating it risks overwriting previous velocity mappings. To this end, we propose OrthoSkillVLA, a parameter-efficient framework for continual skill learning in pretrained VLA models without demonstration replay. Given the representation heterogeneity, we impose separate subspace constraints on the VLM and ActionHead, preserving reusable semantic capacity while protecting localized velocity patterns. For the output layer, we introduce a lightweight feature-aware MoE decoder, where each skill is allocated a compact expert and a training-free router selects the expert according to feature-space affinity. Extensive simulated and real-world evaluations, together with ablations, demonstrate that OrthoSkillVLA better preserves prior skills while acquiring new ones.
发表机构
- School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
- School of Electronic Science and Engineering, Southeast University(东南大学电子科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。