Green-VLA: 通用机器人中的分阶段视觉-语言-动作模型
Green-VLA: Staged Vision-Language-Action Model for Generalist Robots
浏览论文内容
中文总结 AI 辅助
Green-VLA是一种面向真实环境的分阶段视觉-语言-动作模型,通过多阶段训练和强化学习对齐,提升机器人在多样化形态上的泛化能力和性能。
中文摘要 AI 辅助
我们介绍了Green-VLA,一种面向现实世界部署的分阶段视觉-语言-动作(VLA)框架,用于Green人形机器人,同时保持在多样化形态上的泛化能力。Green-VLA遵循五个阶段的课程:(L0)基础VLMs,(L1)多模态接地,(R0)多形态预训练,(R1)形态特定适应,以及(R2)强化学习(RL)策略对齐。我们结合可扩展的数据处理管道(3000小时的演示)与时间对齐和质量过滤,并使用统一的、具有形态意识的动作接口,使单个策略能够控制人形机器人、移动机械臂和固定基臂。在推理过程中,VLA控制器通过episode-progress预测、out-of-distribution检测和基于关节预测的引导来提高安全性和精确目标选择。在Simpler BRIDGE WidowX和CALVIN ABC-D以及真实机器人评估中的实验显示,具有RL对齐的性能在成功率、鲁棒性和长周期效率方面均有显著提升。
英文摘要
We introduce Green-VLA, a staged Vision-Language-Action (VLA) framework for real-world deployment on the Green humanoid robot while maintaining generalization across diverse embodiments. Green-VLA follows a five stage curriculum: (L0) foundational VLMs, (L1) multimodal grounding, (R0) multi-embodiment pretraining, (R1) embodiment-specific adaptation, and (R2) reinforcement-learning (RL) policy alignment. We couple a scalable data-processing pipeline (3,000 hours of demonstrations) with temporal alignment and quality filtering, and use a unified, embodiment-aware action interface enabling a single policy to control humanoids, mobile manipulators, and fixed-base arms. At inference, the VLA controller is enhanced with episode-progress prediction, out-of-distribution detection, and joint-prediction-based guidance to improve safety and precise target selection. Experiments on Simpler BRIDGE WidowX and CALVIN ABC-D, as well as real-robot evaluations, show strong generalization and performance gains from RL alignment in success rate, robustness, and long-horizon efficiency.