发表机构
Peking University; Nanjing University; Simon Fraser University; Hong Kong University of Science and Technology(北京大学; 南京大学; 西蒙弗雷泽大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RoboFL通过MoSAIC联邦专家组装和FARD/PCEA机制,解决世界行动模型数据稀缺与异构问题,在Franka手臂上超越集中式基线12.23%,并减少86.81%通信。
AI 中文摘要
视觉-语言-行动模型和世界行动模型日益流行,但仍受限于稀缺、机构间隔离且任务异质的物理交互数据。一种自然的联邦解决方案是让每个客户端通过参数高效微调来适配共享基础模型,避免交换完整模型更新。然而,联邦化这些适配器并非易事,因为朴素聚合可能纠缠不兼容的更新,而将MoE风格的路由引入联邦聚合可能会稀释专化并破坏专家选择的稳定性。我们提出RoboFL,它实例化了MoSAIC(槽位适配器混合体)用于联邦世界行动学习。MoSAIC直接将本地训练的LoRA适配器作为服务器MoE的专家分支安装。服务器端路由器在这些先验信息分支上学习令牌分配,同时联合优化路由和专家参数。前瞻到行动路由蒸馏(FARD)对齐模型三条路径上的路由,而路径共识专家聚合(PCEA)将完整的专家更新转换为紧凑的全局适配器,用于个性化再分配。在RoboTwin 2.0、RLBench和真实世界Franka机器人手臂上的实验表明,RoboFL凭借结构化专家组装具有优越性,它在Franka手臂上比集中式PEFT InternVLA-A1高出12.23%,同时相对于基于MoE的联邦VLA基线,每轮客户端通信减少高达86.81%。
英文摘要
Vision-language-action and world-action models are increasingly popular, yet remain bottlenecked by physical interaction data that is scarce, institutionally siloed, and task-heterogeneous. A natural federated solution is to let each client adapt a shared foundation model through parameter-efficient fine-tuning, avoiding the exchange of full-model updates. However, federating these adapters is nontrivial, as naive aggregation can entangle incompatible updates, while incorporating MoE-style routing into federated aggregation may dilute specialization and destabilize expert selection. We present RoboFL, which instantiates MoSAIC (Mixture of Slotted Adapters) for federated world-action learning. MoSAIC directly installs locally trained LoRA adapters as the expert branches of a server MoE. Server-side routers learn token assignments over these prior-informed branches while jointly refining routing and expert parameters. Foresight-to-Action Routing Distillation (FARD) aligns routing across the model's three paths, while Path-Consensus Expert Aggregation (PCEA) converts complete expert updates into a compact global adapter for personalized redistribution. Experiments on RoboTwin 2.0, RLBench, and a real-world Franka robot arm show the superiority of RoboFL with structured expert assembly, as it outperforms centralized PEFT InternVLA-A1 by 12.23% on the Franka arm, while reducing per-round client communication by up to 86.81% relative to MoE-based federated VLA baselines.