arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

协作多智能体视觉-语言-动作模型:基于强化微调

Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning

Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo, Hanxiao Li, Kejian Shi, Shuning Zhang, Pu Feng, Yongjia Ma, Yuqing Ma, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Simin Li

arXiv 2609.36588首次发表:更新:

发表机构

Beihang University; The Chinese University of Hong Kong; PKU-Psibot Lab; Tsinghua University; Zhongguancun Laboratory; Li Auto Inc.; Peking University(北京航空航天大学; 香港中文大学; 北大-鹏博实验室; 清华大学; 中关村实验室; 理想汽车公司; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出三阶段强化微调流程,通过初始化感知数据收集、离线信用过滤和在线潜空间微调,提升多智能体VLA的协作能力,在多个任务上显著提高成功率。

AI 中文摘要

我们研究用于协作多智能体视觉-语言-动作(VLA)模型的强化学习(RL)方法。该问题具有挑战性,因为VLA是在大规模单智能体数据上预训练的,因此缺乏机器人间协作所需的细粒度协调技能。在多机器人演示上的监督微调(SFT)部分弥补了这一差距,但其性能受限于演示数据,且无法从自身经验中改进。我们提出了一种用于多智能体VLA的三阶段强化微调(RFT)流程。首先,初始化感知的数据收集扫描初始配置,并仅在预训练VLA反复失败时调用人类演示,从而在降低人力成本的同时增强对初始化偏移的鲁棒性。其次,离线信用过滤调优将信用分配给各个智能体,并在具有正优势的每智能体轨迹上进行微调,而非在整个联合轨迹上微调。第三,我们发现现有的在线RL方法对VLA在困难多智能体任务上效果不佳,这归因于噪声协同探索和不稳定的更新。我们转而使用在线潜空间微调,冻结VLA并在其潜噪声空间中进行RL。我们使用π0和π0.5骨干网络,在RoboTwin、RoboFactory以及两个Franka机器人的真实世界操作中的11个任务上评估了我们的多智能体VLA。我们的多智能体VLA在RoboTwin、RoboFactory和真实世界任务上的平均成功率分别提高了+23.1%、+16.4%和+44%。代码可在该https URL获取。

英文摘要

We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both $π_0$ and $π_{0.5}$ backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by $+23.1\%$, $+16.4\%$, and $+44\%$ on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at https://anonymous.4open.science/r/mavla_rft-2BC0/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑