发表机构
University College London; Cisco Research(伦敦大学学院; 思科研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统评估了三种预训练VLA模型在模拟和真实机器人上的联邦微调,发现联邦范围比聚合算法更关键,并开源了decentvla测试平台以促进公平比较。
AI 中文摘要
将预训练的视觉-语言-动作(VLA)模型适配到新的机器人、环境或任务中,需要本地收集且通常被丢弃的演示数据。联邦学习是一种利用此类分布式演示数据来学习共享策略的有前景的方法。然而,联邦学习能否适配大型预训练VLA模型仍是一个开放问题,并且缺乏针对预训练VLA的可复现基准和可复用的训练框架,使得现有结果难以比较。在本文中,我们对三种现代预训练VLA策略在LIBERO操作基准的40个模拟任务以及两个真实机器人实验中的六个真实世界任务上进行了系统的联邦微调研究,演示数据分别来自两个和三个站点。我们的研究分析了该设置中的关键选择,涵盖多个联邦参数范围、三种聚合算法以及分布偏移下的评估。基于该研究,我们得出了一系列经验教训,包括联邦范围相对于聚合算法选择的优势,以及在物理机器人上匹配集中式微调的难度,其中跨站点异质性比模拟所捕获的更强。我们还强调了联邦VLA学习的机遇,例如能够在异构数据上匹配集中式微调,在分布偏移下至少保持与集中式微调相当的鲁棒性,以及进行个性化,每个客户端对策略的一部分进行联邦,其余部分保持本地,这在策略预训练较弱时有所帮助,但不会留下可用的全局模型。我们开源了decentvla{},这是该研究背后的模型和运行时无关的测试平台,以促进联邦VLA学习的未来研究和公平比较。
英文摘要
Adapting a pretrained Vision-Language-Action (VLA) model to a new robot, environment, or task requires demonstrations that are collected locally and often discarded. Federated learning is a promising approach to exploiting such distributed demonstrations by learning a shared policy. However, whether it can adapt large pretrained VLAs remains an open question, and a lack of reproducible benchmarks for pretrained VLAs and reusable training frameworks makes existing results difficult to compare. In this paper, we conduct a systematic study of federated fine-tuning of three modern pretrained VLA policies on the 40 simulated tasks of the LIBERO manipulation benchmark, and on six real-world tasks in two real-robot experiments, with demonstrations collected across two and three sites, respectively. Our study analyzes the key choices in this setting, spanning multiple federated parameter scopes, three aggregation algorithms, and evaluation under distribution shift. Based on the study, we derive a series of lessons, including the dominance of the federated scope over the choice of aggregation algorithm and the difficulty of matching centralized fine-tuning on physical robots, where cross-site heterogeneity is stronger than simulation captures. We also highlight opportunities for federated VLA learning, such as the ability to match centralized fine-tuning on heterogeneous data, to remain at least as robust as centralized fine-tuning under distribution shift, and to personalize, with each client federating part of the policy and keeping the rest local, which helps where the policy's pretraining is weak but leaves no usable global model. We open-source \decentvla{}, the model- and runtime-agnostic testbed behind the study, to facilitate future research and fair comparisons in federated VLA learning.