发表机构
Li Auto Inc.; School of Artificial Intelligence, Beijing University of Posts and Telecommunications; The Chinese University of Hong Kong, Shenzhen(理想汽车; 北京邮电大学人工智能学院; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出VLAFlow统一框架,比较四种训练范式,发现语言监督保持视觉-语言泛化,未来潜在对齐改进状态转换建模,两者结合实现最佳迁移性能。
AI 中文摘要
视觉-语言-动作模型(VLA)最近推动了机器人操作的发展,然而不同机器人数据预训练范式的影响仍然难以比较,因为现有模型在架构、数据、动作空间和评估协议上往往不同。我们提出VLAFlow(视觉-语言-动作流),一个用于VLA训练目标受控比较的统一流匹配框架。使用包含来自DROID、OpenX-Embodiment、OpenX-Augmented和RoboCOIN的约5000小时数据的异构机器人语料库OXEMix,我们在相同的pi0风格架构、共享VLM骨干、动作专家和14维动作空间下评估了四种范式:纯动作建模(MindPI)、语言监督协同训练(MindLPI)、未来潜在对齐(MindWPI)及其组合(MindLWPI)。在LIBERO、LIBERO-Plus和SimplerEnv上的实验表明,纯动作预训练对异构数据敏感。相比之下,语言监督有助于保持视觉-语言泛化,而未来潜在对齐改进了状态转换和动作结果建模。通过结合这两种信号,MindLWPI在基准测试中实现了最稳定的整体迁移性能。这些结果表明了一种元动作空间观点:语言和未来潜在表示提供了互补的中间约束,使得异构动作监督更加平滑和可迁移。
英文摘要
Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training paradigms remain difficult to compare because existing models often differ in architecture, data, action space, and evaluation protocol. We present VLAFlow (Vision-Language-Action Flow), a unified flow-matching framework for controlled comparison of VLA training objectives. Using a heterogeneous robot corpus, OXEMix, containing approximately 5,000 hours of data from DROID, OpenX-Embodiment, OpenX-Augmented, and RoboCOIN, we evaluate four paradigms under the same pi0-style architecture, shared VLM backbone, action expert, and 14-dimensional action space: action-only modeling (MindPI), language-supervised co-training (MindLPI), future latent alignment (MindWPI), and their combination (MindLWPI). Experiments on LIBERO, LIBERO-Plus, and SimplerEnv show that action-only pre-training is sensitive to heterogeneous data. In contrast, language supervision helps preserve vision-language generalization, while future latent alignment improves state-transition and action-outcome modeling. By combining both signals, MindLWPI achieves the most stable overall transfer performance across benchmarks. These results suggest a meta-action space view: language and future latent representations provide complementary intermediate constraints that make heterogeneous action supervision smoother and more transferable.