arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AirMoE:在无线边缘实现无线分布式混合专家(Mixture-of-Experts,MoE)推理

AirMoE: Realizing Over-the-Air Distributed Mixture-of-Experts Inference at the Wireless Edge

Huiling Yang, Zhanwei Wang, Kaibin Huang

arXiv 2608.22932首次发表:更新:

AI 中文总结

AirMoE是利用AirComp实现无线分布式MoE推理的框架,通过双时间尺度策略解决AirComp带来的挑战,在ARC-Easy基准上显著提升了E2E推理准确性,尤其适用于强设备异构场景。

AI 中文摘要

混合专家(MoE)架构通过稀疏的专家激活减少每令牌计算,从而支持无线边缘的高效大语言模型(LLM)推理。无线分布式MoE(WIDE)架构通过由边缘服务器协调,将计算密集型的专家分布在不同设备上,以解决边缘设备的资源约束。然而,通过正交多址重复上传高维专家输出会产生严重的上行链路瓶颈。为克服这一限制,我们提出AirMoE,这是一种利用无线计算(AirComp)实现无线波形叠加的同时专家输出聚合的新型框架。然而,将AirComp集成到MoE推理中会带来三个独特挑战:快速变化的聚合权重、依赖层的误差敏感性以及感知信道的专家放置。为解决这些挑战,我们首先构建推理感知的AirMoE误差指标,通过基于扰动的层敏感性校准,量化物理层聚合失真对端到端(E2E)推理准确性的影响。然后,我们将最小化该E2E误差的联合优化问题公式化,并在不损失最优性的情况下将其分解为双时间尺度框架。在快时间尺度上,我们推导了全局最优的基于阈值的功率控制策略,该策略将设备划分为两组:一组实现精确的聚合权重对齐,其余设备以最大功率传输。在慢时间尺度上,我们开发了感知激活和信道的专家放置策略,将更重要的专家分配给具有更低信道功率成本的设备。使用ARC-Easy基准上的OLMoE-1B-7B-0924模型进行的大量实验表明,AirMoE在E2E推理准确性上显著优于代表性基线,特别是在设备异构性强的情况下。

英文摘要

Mixture-of-experts (MoE) architectures enable efficient large language model (LLM) inference at the wireless edge through sparse activation. The wireless distributed MoE (WIDE) architecture addresses edge-resource constraints by distributing experts across devices coordinated by an edge server. However, WIDE suffers from repeated uplink transmissions of high-dimensional expert outputs via orthogonal multiple access. To overcome this bottleneck, we propose AirMoE, an over-the-air computing (AirComp)-enabled framework for simultaneous expert-output aggregation via wireless waveform superposition. Integrating AirComp into MoE inference introduces three challenges: fast-varying aggregation weights, layer-dependent error sensitivity, and channel-aware expert placement. To address these challenges, we construct an inference-aware AirMoE error metric to quantify aggregation distortion effects on end-to-end (E2E) inference accuracy via perturbation-based layer-sensitivity calibration. We then formulate a joint optimization problem to minimize this error and decompose it, without loss of optimality, into a two-timescale framework. At the fast timescale, we derive a globally optimal threshold-based power-control policy that separates devices into coefficient-aligned and full-power groups. At the slow timescale, we develop an activation- and channel-aware expert placement strategy that assigns more important experts to devices with lower channel-power cost. Extensive experiments demonstrate that AirMoE outperforms representative baselines in E2E inference accuracy, especially under strong device heterogeneity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑