Co-VLA:基于共识的视觉-语言-动作模型联邦训练
Co-VLA: Consensus-based Federated Training for Vision-Language-Action Models
浏览论文内容
中文总结 AI 辅助
Co-VLA提出基于ADMM共识优化的联邦训练方法,用于视觉-语言-动作模型,支持全模型训练和参数高效微调,在分散机器人数据上达到与集中训练相当的性能。
中文摘要 AI 辅助
视觉-语言-动作模型(VLAs)已成为通用机器人学习的一种有前景的范式,其性能随着模型和数据集的扩展而提升。然而,扩展机器人数据收集仍然具有挑战性,因为数据自然分布在不同的机器人、任务和地点之间,使得集中化成本高昂或不切实际。联邦学习提供了一种在分散的机器人数据上进行训练的方法,但将其应用于VLAs需要考虑到异构的机器人客户端数据分布。我们提出了Co-VLA,它利用交替方向乘子法(ADMM)将共识优化应用于联邦VLA训练。我们展示了同一算法同时支持全模型训练以及使用固定秩和秩自适应适配器的参数高效微调。Co-VLA这个名字既体现了共识也体现了协作:拥有不同本地机器人数据集的客户端协作训练一个共享模型,而无需共享它们的数据。我们的实验表明,在全模型训练和参数高效微调设置中,Co-VLA达到了与集中式训练相当的性能。
英文摘要
Vision-language-action models (VLAs) have emerged as a promising paradigm for general-purpose robot learning, with performance improving as models and datasets scale. Scaling robot data collection, however, remains challenging because data are naturally distributed across robots, tasks, and locations, making centralization costly or impractical. Federated learning offers a way to train on decentralized robot data, but applying it to VLAs requires accounting for heterogeneous robot client data distributions. We present Co-VLA, which applies consensus optimization using the Alternating Direction Method of Multipliers~(ADMM) to federated VLA training. We show that the same algorithm supports both full-model training and parameter-efficient fine-tuning with both fixed-rank and rank-adaptive adapters. The name Co-VLA reflects both consensus and collaboration: clients with different local robot datasets collaboratively train a shared model without sharing their data. Our experiments demonstrate that Co-VLA achieves performance comparable to centralized training in both full-model training and parameter-efficient fine-tuning settings. The project website and videos of our real-world experiments are available at https://embodiedvision.github.io/co-vla/.
发表机构
- University of Augsburg(奥格斯堡大学)
- Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
机构由 AI 辅助整理,请以论文原文为准。