一起训练还是之后合并?通过共享动作接口统一VLA专家
Train Together or Merge Later? Unifying VLA Experts via a Shared Action Interface
AI总结:
提出PolicyWeave,通过保留共享动作接口和上下文引导的稀疏合并,统一独立训练的VLA专家,在RoboCasa365上显著提升多任务成功率。
AI中文摘要:
共同训练提供了一种构建多任务视觉-语言-动作(VLA)策略的直接方式,但可能达不到独立训练每个任务所实现的性能。挑战在于,在没有联合后训练的情况下,在多任务策略中保留这些任务特定的增益。通过模型合并组合独立训练的专家是一种自然的方法,然而强大的个体专家并不一定产生强大的合并策略。我们识别了这种不兼容的一个来源:动作接口的任务特定更改,包括动作归一化以及动作编码器和解码器。我们提出PolicyWeave,将兼容合并的后训练与上下文引导的稀疏合并相结合。在后训练期间,所有专家保留共同基础策略的动作接口,而任务适应仅限于动作模型隐藏层中的LoRA更新。这使得专家与现有的模型合并方法更加兼容。然而,合并所有专家在部署时仍可能引入来自无关任务的干扰。PolicyWeave使用初始视觉-语言上下文对每个专家的LoRA更新进行评分,通过留一层排序稳定性确定专家集,并形成所选更新的稀疏加权合并,该合并对于当前任务保持固定。我们使用GR00T N1.5在18个RoboCasa365任务上评估PolicyWeave,仅使用目标任务演示的10%进行监督微调(SFT)。保留共享动作接口将四种静态合并方法的平均成功率从17.0%提高到52.8%。PolicyWeave使用这些SFT专家达到64.7%的成功率,并在任务特定强化学习(RL)后达到74.1%,而联合RL为60.7%。在LIBERO-10和AgileX Piper机械臂上的进一步评估支持在长时程和真实世界操作中部署独立学习的技能。
英文摘要:
Co-training offers a straightforward way to build a multi-task vision-language-action (VLA) policy, but can fall short of the performance achieved by training each task independently. The challenge is to retain these task-specific gains in a multi-task policy without joint post-training. Combining independently trained experts through model merging is a natural approach, yet strong individual experts do not necessarily yield a strong merged policy. We identify one source of this incompatibility: task-specific changes to the action interface, comprising action normalization and the action encoder and decoder. We propose PolicyWeave, combining merge-compatible post-training with context-guided sparse merging. During post-training, all experts retain the common base policy's action interface, while task adaptation is restricted to LoRA updates in the hidden layers of the action model. This makes the experts more compatible with existing model merging methods. However, merging all experts can still introduce interference from unrelated tasks at deployment. PolicyWeave scores each expert's LoRA updates using the initial visual-language context, determines the expert set through leave-one-layer-out ranking stability, and forms a sparse weighted merge of the selected updates that remains fixed for current task. We evaluate PolicyWeave with GR00T N1.5 on 18 RoboCasa365 tasks, using only 10% of the target-task demonstrations for supervised fine-tuning (SFT). Preserving the shared action interface raises the average success rate across four static merging methods from 17.0% to 52.8%. PolicyWeave achieves 64.7% success with these SFT experts and 74.1% after task-specific reinforcement learning (RL), compared with 60.7% for joint RL. Further evaluations on LIBERO-10 and an AgileX Piper arm support the deployment of independently learned skills in long-horizon and real-world manipulation.