X-VLA:作为可扩展跨本体视觉-语言-动作模型的软提示 Transformer
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
浏览论文内容
中文总结 AI 辅助
提出基于流匹配的跨本体视觉-语言-动作模型 X-VLA,通过引入本体特定的软提示嵌入有效利用异构机器人数据,在多项模拟和真实机器人基准上实现 SOTA 性能。
中文摘要 AI 辅助
成功的通用视觉-语言-动作(VLA)模型依赖于在不同机器人平台上进行有效训练,并使用大规模、跨本体、异构数据集。为了促进和利用丰富、多样化机器人数据源中的异构性,我们提出了一种新颖的软提示方法,通过将提示学习概念引入跨本体机器人学习,并为每个不同的数据源引入独立的可学习嵌入集合,从而最小化新增参数。这些嵌入作为本体特定的提示,共同赋予 VLA 模型有效利用不同跨本体特征的能力。我们的新模型 X-VLA 是一种基于流匹配的简洁 VLA 架构,完全依赖于软提示的标准 Transformer 编码器,兼具可扩展性与简单性。在 6 个模拟环境和 3 个真实世界机器人上进行评估,我们的 0.9B 实例化版本 X-VLA-0.9B 在一系列基准测试中同时实现了 SOTA 性能,在从灵活灵巧性到跨本体、环境和任务的快速适应等广泛能力维度上展现出卓越结果。
英文摘要
Successful generalist Vision-Language-Action (VLA) models rely on effective training across diverse robotic platforms with large-scale, cross-embodiment, heterogeneous datasets. To facilitate and leverage the heterogeneity in rich, diverse robotic data sources, we propose a novel Soft Prompt approach with minimally added parameters, by infusing prompt learning concepts into cross-embodiment robot learning and introducing separate sets of learnable embeddings for each distinct data source. These embeddings serve as embodiment-specific prompts, which in unity empower VLA models with effective exploitation of varying cross-embodiment features. Our new X-VLA, a neat flow-matching-based VLA architecture, relies exclusively on soft-prompted standard Transformer encoders, enjoying both scalability and simplicity. Evaluated across 6 simulations as well as 3 real-world robots, our 0.9B instantiation-X-VLA-0.9B simultaneously achieves SOTA performance over a sweep of benchmarks, demonstrating superior results on a wide axes of capabilities, from flexible dexterity to quick adaptation across embodiments, environments, and tasks. Website: https://thu-air-dream.github.io/X-VLA/
发表机构
- Institute for AI Industry Research (AIR) Tsinghua University(人工智能产业研究院(AIR)清华大学)
- Shanghai AI Lab(上海人工智能实验室)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。