arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AirMoE:用于协同智能的统计增强无线空中MoE

AirMoE: Statistic-Augmented Over-the-Air MoE for Collaborative Intelligence

Wei-Bin Kou, Jingreng Lei, Guangxu Zhu, Yujiu Yang

arXiv 2607.16562首次发表:更新:

AI 中文总结

研究在无线云边缘网络部署混合专家(MoE)的瓶颈问题,提出统计增强无线空中MoE(AirMoE)范式,包括路由和聚合两方面机制,经理论分析与实验验证,该范式优于基线和单模型,证实各组件有效性。

AI 中文摘要

混合专家(MoE)越来越多地部署在无线云边缘网络上,因为单个边缘设备在本地缺乏承载大规模模型的足够资源。在这种分布式架构中,云托管的预训练大模型(LM)作为潜在特征提取的共享主干,而异质专家分布在分布式无线连接的客户端上协同形成任务头。然而,在无线链路上部署MoE存在两个耦合瓶颈。一方面,由于需要传输原始特征,路由激活哪些客户端通常会使带宽受限的上行链路过载。另一方面,通过无线链路聚合激活专家的输出受到信道噪声和扩展性差的阻碍。为了打破这些瓶颈,我们提出了一种统计增强的无线空中MoE(AirMoE)范式。具体来说,在路由方面,每个客户端通过云广播的紧凑查询查询其本地特征检索库(FRL),检索原型诱导统计量,并将其数字报告给云,大大减少上行链路流量;云然后通过Jensen-Shannon(JS)散度将这些统计量与LM提取的特征对齐来选择最相关的客户端。在聚合方面,选定的专家通过多址信道同时传输他们的输出,通过波形叠加物理计算重新加权的和,通过信道感知功率控制实现重新加权系数。这两种机制在算法和物理上都是解耦的。我们进一步提供了关于收敛和迭代复杂度的理论分析。以语义分割任务为例,大量实验表明AirMoE优于MoE基线和单模型竞争对手。消融实验进一步证实了每个纳入组件的有效性。

英文摘要

Mixture of Experts (MoE) are increasingly deployed over wireless cloud-edge networks, as a single edge device lacks sufficient resources to host large-scale models locally. In this distributed architecture, a cloud-hosted pretrained Large Model (LM) acts as a shared backbone for latent feature extraction, while heterogeneous experts deployed across distributed, wirelessly-connected clients collaboratively form the task head. However, deploying MoE over wireless links exposes two coupled bottlenecks. On the one hand, routing which clients to activate generally overloads bandwidth-limited uplinks due to required raw feature transmission. On the other hand, aggregating the activated experts' outputs over wireless links is hindered by channel noise and poor scalability. To break these bottlenecks, we propose a statistic-augmented over-the-air MoE (AirMoE) paradigm. Specifically, on the routing side, each client queries its local Feature Retrieval Library (FRL) with a cloud-broadcast compact query, retrieves a prototype-induced statistic, and reports it digitally to the cloud, drastically reducing uplink traffic; the cloud then selects the most relevant clients by aligning these statistics with the LM-extracted features via Jensen--Shannon (JS) divergence. On the aggregating side, selected experts simultaneously transmit their outputs over the multiple-access channel, which physically computes the reweighted sum via waveform superposition, with reweighting coefficients realized through channel-aware power control. The two mechanisms are thus decoupled both algorithmically and physically. We further provide theoretical analyses on convergence and iteration complexity. Taking semantic segmentation task as an example, extensive experiments demonstrate that AirMoE outperforms MoE baselines and single-model competitors. Ablations further confirm the effectiveness of each incorporated component.

Comments17 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑