arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

联邦子空间引导的视觉-语言-动作策略蒸馏用于非独立同分布多机器人操作

Federated Subspace Guided Vision-Language-Action Policy Distillation for Non-IID Multi-Robot Manipulation

Biprodip Pal, Kaushik Roy, Yanming Zhu, Brendan Tidd, Alan Wee-Chung Liew, Peyman Moghadam

arXiv 2609.32239首次发表:更新:

AI 中文总结

提出FedDRMan联邦子空间引导蒸馏框架,通过低秩多模态子空间蒸馏和基于更新兼容性的聚类聚合,解决非独立同分布多机器人操作中的表征漂移,在LIBERO上以80.7%成功率超越强基线11.6个百分点。

AI 中文摘要

联邦学习为多个机器人提供了一种自然的方式,在不集中访问训练示范的情况下共同改进操作策略。然而,非独立同分布的任务和环境分布可能导致表征漂移和相互不兼容的机器人策略更新,使得朴素的参数聚合具有破坏性。我们提出了FedDRMan,一个用于异构机器人操作的联邦子空间引导蒸馏框架。在每个通信轮次中,服务器模型提供一个冻结的教师模型用于本地行为克隆,而低秩多模态子空间和动作分布蒸馏则保留全局有用的表征几何和策略行为。为了解决异构聚合问题,FedDRMan根据更新兼容性对客户端进行分组,并为每个簇维护一个持久模型。然后,服务器对每个兼容的聚合进行谱再平衡,以减轻较弱的任务相关机器人策略更新方向的衰减。在LIBERO上进行的广泛实验,涵盖多样的非独立同分布设置、异构性水平、客户端参与变化,以及消融和聚合分析,表明FedDRMan显著提高了知识迁移,并持续优于强大的联邦基线,达到了80.7%的峰值平均成功率,比评估的最强联邦基线高出11.6个百分点。

英文摘要

Federated learning offers a natural way for multiple robots to jointly improve manipulation policies without requiring centralized access to training demonstrations. However, non-IID task and environment distributions can induce representation drift and mutually incompatible robot-policy updates, making naive parameter aggregation destructive. We present FedDRMan, a federated subspace-guided distillation framework for heterogeneous robot manipulation. At each communication round, the server model provides a frozen teacher for local behavior cloning, while low-rank multimodal subspace and action-distribution distillation preserve globally useful representation geometry and policy behavior. To address heterogeneous aggregation, FedDRMan groups clients by update compatibility and maintains a persistent model for each cluster. The server then spectrally rebalances each compatible aggregate to mitigate attenuation of weaker task-relevant robot-policy update directions. Extensive experiments on LIBERO across diverse non-IID settings, heterogeneity levels, client participation variation, together with ablations and aggregation analyses, show that FedDRMan substantially improves knowledge transfer and consistently outperforms strong federated baselines achieving a peak mean success rate of 80.7%, 11.6 percentage points above the strongest evaluated federated baseline.

Comments9 Pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑