LAP:语言-动作预训练实现零样本跨躯体迁移
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
- Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
LAP通过语言-动作预训练实现零样本跨躯体迁移,首次在无需特定躯体微调的情况下达到显著的迁移效果,提升性能达两倍。
AI中文摘要:
机器人领域长期以来的目标是开发一种通用策略,能够在无需针对新机器人躯体进行适应的情况下实现零样本部署。尽管存在大规模多躯体预训练,现有的视觉-语言-动作模型(VLAs)仍紧密耦合于其训练躯体,并且通常需要成本高昂的微调。我们引入了语言-动作预训练(LAP),这是一种简单的配方,直接将低层机器人动作表示为自然语言,使动作监督与预训练视觉-语言模型的输入-输出分布对齐。LAP不需要学习的分词器,不需要成本高昂的标注,也不需要特定躯体的架构设计。基于LAP,我们提出了LAP-3B,据我们所知,这是首个在无需任何特定躯体微调的情况下实现显著零样本迁移的VLA。在多个新型机器人和操作任务中,LAP-3B实现了超过50%的平均零样本成功率,比最强的先前VLAs提高了约两倍。我们进一步表明,LAP能够实现高效的适应和良好的扩展性,同时通过共享的语言-动作格式统一了动作预测和VQA,并通过联合训练获得额外的收益。
英文摘要:
A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision-Language-Action models (VLAs) remain tightly coupled to their training embodiments and typically require costly fine-tuning. We introduce Language-Action Pre-training (LAP), a simple recipe that represents low-level robot actions directly in natural language, aligning action supervision with the pre-trained vision-language model's input-output distribution. LAP requires no learned tokenizer, no costly annotation, and no embodiment-specific architectural design. Based on LAP, we present LAP-3B, which to the best of our knowledge is the first VLA to achieve substantial zero-shot transfer to previously unseen robot embodiments without any embodiment-specific fine-tuning. Across multiple novel robots and manipulation tasks, LAP-3B attains over 50% average zero-shot success, delivering roughly a 2x improvement over the strongest prior VLAs. We further show that LAP enables efficient adaptation and favorable scaling, while unifying action prediction and VQA in a shared language-action format that yields additional gains through co-training.