MoS-VLA:一种具备单次技能适应能力的视觉-语言-动作模型
MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
- University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出 MoS-VLA 框架,通过跨 Open X-Embodiment 数据集学习技能基函数,并利用一次演示和凸优化组合新技能,实现无需梯度更新的快速机器人策略适应。
AI中文摘要:
在大型机器人数据集上训练的视觉-语言-动作(Vision-Language-Action,VLA)模型,有望在不同领域和实体形态之间实现通用、鲁棒的控制。然而,现有方法在部署到新环境、新实体形态或新任务时,往往无法即开即用。我们提出 Mixture of Skills VLA(MoS-VLA)框架,它将机器人操作策略表示为有限的一组已学习基函数的线性组合。在预训练阶段,MoS-VLA 利用来自 Open X-Embodiment 项目的多个数据集联合学习这些基函数,从而形成结构化的技能空间。在测试时,适应一项新任务只需要一次专家演示。随后,通过一个最小化 L1 动作误差的轻量级凸优化问题推断相应的技能表示,无需梯度更新。这种无梯度适应只会产生极小的额外开销,同时能够快速实例化新技能。实验结果表明,MoS-VLA 在全部五个未见过的数据集上均取得更低的动作预测误差,并且在仿真和真实机器人任务中均取得成功,而预训练 VLA 模型在这些任务中会彻底失败。项目页面:mos-vla.github.io/
英文摘要:
Vision-Language-Action (VLA) models trained on large robot datasets promise general-purpose, robust control across diverse domains and embodiments. However, existing approaches often fail out-of-the-box when deployed in novel environments, embodiments, or tasks. We introduce Mixture of Skills VLA (MoS-VLA), a framework that represents robot manipulation policies as linear combinations of a finite set of learned basis functions. During pretraining, MoS-VLA jointly learns these basis functions across datasets from the Open X-Embodiment project, producing a structured skill space. At test time, adapting to a new task requires only a single expert demonstration. The corresponding skill representation is then inferred via a lightweight convex optimization problem that minimizes the L1 action error, without requiring gradient updates. This gradient-free adaptation incurs minimal overhead while enabling rapid instantiation of new skills. Empirically, MoS-VLA achieves lower action-prediction error on five out of five unseen datasets and succeeds in both simulation and real-robot tasks where a pretrained VLA model fails outright. Project page: mos-vla.github.io/