发表机构
The Hong Kong University of Science and Technology (Guangzhou); Shanghai Jiao Tong University; Eastern Institute of Technology; ITMO University; University of Hong Kong; Singapore University of Technology and Design; Mohamed Bin Zayed University of Artificial Intelligence(香港科技大学(广州); 上海交通大学; 东方理工大学; 圣彼得堡国立信息技术机械与光学大学; 香港大学; 新加坡科技设计大学; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对6G具身智能中VLA模型联邦训练面临隐私、通信和异构性挑战,提出模态解耦联邦学习框架FedMVLA,通过模态感知聚合、隐私分配和通信压缩,在干扰下实现84.8%任务成功率并大幅降低通信负载。
AI 中文摘要
第六代(6G)无线网络有望为大规模具身智能提供关键基础设施,其中异构机器人通过低延迟连接、边缘智能和分布式感知进行协作。视觉-语言-动作(VLA)模型通过将视觉感知、语言理解和动作生成整合为统一的闭环策略,为此提供了基础。然而,将VLA模型训练并适配到分布式机器人智能体上,带来了隐私保护、通信效率和模型异构性方面的挑战。现有的联邦学习(FL)方法忽视了视觉、语言和动作通路在参数规模、隐私暴露、更新动态以及对压缩或扰动的容忍度方面的内在差异。为解决这一问题,本文提出了FedMVLA,一种面向6G网络隐私保护具身智能的模态解耦联邦学习框架。FedMVLA包含三种机制:模态感知联邦聚合(MAFA)、模态感知隐私分配(MAPA)和模态感知通信压缩(MACO),并配合一种模态切片传输设计,将精度关键的动作流通过受保护的超可靠低延迟切片进行路由。一项基于第三代合作伙伴计划(3GPP)无线底层(涵盖衰落、同信道干扰和恶意干扰)的联邦机器人操作案例研究表明,FedMVLA实现了84.8%的任务成功率,超过FedAvg 22.2个百分点,在扩展到八个小区中的128个客户端时保持不断扩大的优势,并将调度平均的每客户端上行模型更新负载减少了95.6%(约96%),同时将轮次关键上行完成时间的第95百分位(p95)保持在接近1.5秒。
英文摘要
Sixth-generation (6G) wireless networks are expected to provide a key infrastructure for large-scale embodied intelligence, where heterogeneous robots collaborate through low-latency connectivity, edge intelligence, and distributed sensing. Vision-language-action (VLA) models offer a foundation by integrating visual perception, language understanding, and action generation into a unified closed-loop policy. However, training and adapting VLA models to distributed robotic agents introduce challenges in privacy protection, communication efficiency, and model heterogeneity. Existing federated learning (FL) methods overlook the intrinsic differences among vision, language, and action pathways in parameter scale, privacy exposure, update dynamics, and tolerance to compression or perturbation. To address this issue, this article proposes FedMVLA, a modality-decoupled FL framework for privacy-preserving embodied intelligence in 6G networks. FedMVLA incorporates three mechanisms: modality-aware federated aggregation (MAFA), modality-aware privacy allocation (MAPA), and modality-aware communication compression (MACO), together with a modality-sliced transport design that routes the precision-critical action stream through a protected ultra-reliable low-latency slice. A case study on federated robotic manipulation over the Third Generation Partnership Project (3GPP)-based wireless substrate, covering fading, co-channel interference, and malicious jamming, shows that FedMVLA achieves an 84.8% task success rate, exceeds FedAvg by 22.2 percentage points, sustains a widening margin when scaling to 128 clients across eight cells, and reduces the schedule-averaged per-client uplink model-update payload by 95.6% (approximately 96%), while keeping the 95th percentile (p95) of the round-critical uplink completion time near 1.5s.
CommentsThis article has been accepted for publication in IEEE Wireless Commnunications Magazine