发表机构
Beihang University; The Chinese University of Hong Kong; PKU-Psibot Lab; Tsinghua University; Zhongguancun Laboratory; Li Auto Inc.; Peking University(北京航空航天大学; 香港中文大学; 北大-鹏博实验室; 清华大学; 中关村实验室; 理想汽车公司; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出三阶段强化微调流程,通过初始化感知数据收集、离线信用过滤和在线潜空间微调,提升多智能体VLA的协作能力,在多个任务上显著提高成功率。
AI 中文摘要
我们研究用于协作多智能体视觉-语言-动作(VLA)模型的强化学习(RL)方法。该问题具有挑战性,因为VLA是在大规模单智能体数据上预训练的,因此缺乏机器人间协作所需的细粒度协调技能。在多机器人演示上的监督微调(SFT)部分弥补了这一差距,但其性能受限于演示数据,且无法从自身经验中改进。我们提出了一种用于多智能体VLA的三阶段强化微调(RFT)流程。首先,初始化感知的数据收集扫描初始配置,并仅在预训练VLA反复失败时调用人类演示,从而在降低人力成本的同时增强对初始化偏移的鲁棒性。其次,离线信用过滤调优将信用分配给各个智能体,并在具有正优势的每智能体轨迹上进行微调,而非在整个联合轨迹上微调。第三,我们发现现有的在线RL方法对VLA在困难多智能体任务上效果不佳,这归因于噪声协同探索和不稳定的更新。我们转而使用在线潜空间微调,冻结VLA并在其潜噪声空间中进行RL。我们使用π0和π0.5骨干网络,在RoboTwin、RoboFactory以及两个Franka机器人的真实世界操作中的11个任务上评估了我们的多智能体VLA。我们的多智能体VLA在RoboTwin、RoboFactory和真实世界任务上的平均成功率分别提高了+23.1%、+16.4%和+44%。代码可在该https URL获取。
英文摘要
We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both $π_0$ and $π_{0.5}$ backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by $+23.1\%$, $+16.4\%$, and $+44\%$ on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at https://anonymous.4open.science/r/mavla_rft-2BC0/.