发表机构
Chongqing University; City University of Hong Kong (Dongguan); Shenzhen Future Network of Intelligence Institute; the Chinese University of Hong Kong (Shenzhen)(重庆大学; 香港城市大学(东莞); 深圳未来智能网络研究院; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对无线边缘分布式MoE推理的聚合瓶颈,提出logit感知的MIMO AirComp框架,联合优化接收合并与预编码,在30dB下保留99.3%的GSM8K准确率,显著优于传统AirComp。
AI 中文摘要
分布式混合专家(MoE)推理是一种在无线边缘网络部署大型语言模型(LLM)的有前景架构,因为稀疏专家可以分布在协调的基站(BS)上,而锚节点和用户设备(UE)可以将LLM推理任务卸载到基站。通信瓶颈在于MoE聚合,锚基站必须在每个解码步骤之前恢复所选专家输出的加权和。空中计算(AirComp)与此操作高度匹配,因为无线多址信道自然叠加同时传输。然而,传统AirComp最小化通信失真,而MoE聚合误差对LLM输出的影响不均等。我们提出了一种logit感知的MIMO AirComp框架,将局部logit敏感性估计为块级权重,并在每基站功率约束下联合优化接收合并器和基站预编码器。我们还开发了一种交替算法,保护解码决策免受聚合扰动的影响。使用OpenCompass,我们在GSM8K和ARC-Challenge上评估了Qwen3-30B-A3B-Instruct-2507-FP8。在30 dB聚合信噪比下,受扰动的Qwen3保留了干净GSM8K准确率的99.3%和干净ARC-Challenge准确率的94.8%。在295个样本的直接ARC-Challenge闭环审计中,SW-AirComp在30 dB下达到43.73%的准确率,在20 dB下达到26.78%,优于未加权、匹配滤波和迫零AirComp。在无线模拟器中,SW-AirComp减少了决策相关的聚合失真;在20 dB下,其最终加权和均方误差在相同信道、样本和敏感性权重下比未加权AirComp低53.5%。Logit-RMSE和logit-gap诊断进一步表明,增益来自于减少输出logit扰动和降低top-token变化的风险。
英文摘要
Distributed mixture-of-experts (MoE) inference is a promising architecture for deploying large language models (LLMs) at wireless edge networks because sparse experts can be placed across coordinated base stations (BSs), while the anchor node and user equipment (UE) can offload LLM inference tasks to BSs. The communication bottleneck is the MoE aggregation, where the anchor BS must recover a weighted sum of selected expert outputs before each decoding step. Over-the-air computation (AirComp) is well matched to this operation because the wireless multiple-access channel naturally superposes simultaneous transmissions. However, conventional AirComp minimizes communication distortion, whereas MoE aggregation errors have unequal impact on LLM outputs. We propose a logit-aware MIMO AirComp framework that estimates local logit sensitivity as block-level weights and jointly optimizes receive combiners and BS precoders under per-BS power constraints. We also develop an alternating algorithm that protects decoding decisions from aggregation perturbations. Using OpenCompass, we evaluate Qwen3-30B-A3B-Instruct-2507-FP8 on GSM8K and ARC-Challenge. At 30 dB aggregation SNR, perturbed Qwen3 retains 99.3% of clean GSM8K accuracy and 94.8% of clean ARC-Challenge accuracy. In a direct ARC-Challenge closed-loop audit over 295 examples, SW-AirComp achieves 43.73% accuracy at 30 dB and 26.78% at 20 dB, outperforming unweighted, matched-filter, and zero-forcing AirComp. In the wireless simulator, SW-AirComp reduces decision-relevant aggregation distortion; at 20 dB, its final weighted-sum mean-squared error is 53.5% lower than unweighted AirComp under the same channels, samples, and sensitivity weights. Logit-RMSE and logit-gap diagnostics further show that the gain comes from reducing output-logit perturbation and lowering the risk of top-token changes.
Comments13 pages, 9 figures