Wi-Fi中的广播速率限制:协作式边缘大语言模型(LLM)推理被遗忘的瓶颈
Broadcast Rate Limits in Wi-Fi: A Forgotten Bottleneck for Collaborative Edge LLM Inference
AI总结:
该研究针对协作式边缘LLM推理的通信瓶颈,提出适配UDP广播的MoE推理方法,发现Wi-Fi广播速率限制的遗留问题并通过仿真验证,推动Wi-Fi标准将广播纳入高吞吐量数据平面。
AI中文摘要:
大语言模型(LLM)的部署正从数据中心迁移至边缘设备,混合专家(MoE)模型为此提供了可行路径:稀疏的专家激活机制可让模型分布在多个低成本边缘节点上。分布式MoE推理会反复将嵌入向量从主节点发送至多个工作节点,这种一对多模式难以被主流栈(NCCL、TCP)的单播传输有效支持,却与UDP广播天然适配。我们提出一种用于协作式边缘MoE推理的UDP广播方法,该方法结合了超时驱动重传(利用分布式MoE近乎确定的延迟保障可靠性)与无序结果收集(应对专家预测错误以提升鲁棒性),在8节点有线集群上实现了较NCCL和TCP稳定的1.4倍加速。然而在无线场景中,我们发现了一个更深层、被长期遗忘的瓶颈:IEEE 802.11标准将广播速率上限设为54 Mbps,这一遗留策略原本针对稀疏控制流量设计,并不适配边缘AI需求。在1米、2米和5米距离下开展的NS-3仿真显示,最优速率分别为标准上限54 Mbps的64倍、43倍和32倍。因此,我们认为广播不再是控制平面的遗留产物,Wi-Fi标准应将其视为高吞吐量数据平面的核心组成部分。
英文摘要:
LLM deployment is migrating from data centers to edge devices, where Mixture-of-Experts (MoE) models offer a promising path: sparse expert activation allows the model to be spread across multiple low-cost edge nodes. Distributed MoE inference repeatedly dispatches embeddings from one main node to many workers - a one-to-many pattern poorly served by the sequential unicasts of mainstream stacks (NCCL, TCP), yet naturally matched by UDP broadcast. We propose a UDP broadcast method for collaborative edge MoE inference, augmented with timeout-driven retransmission exploiting near deterministic latency in distributed MoE for reliability and unordered result gathering for robustness to expert mispredictions, yielding a consistent 1.4x speedup over NCCL and TCP on a wired 8-node cluster. In wireless settings, however, we uncover a deeper, long-forgotten bottleneck: IEEE 802.11 caps broadcast rates at 54 Mbps regardless of physical-layer capacity - a legacy policy built for sparse control traffic, not edge AI. NS-3 simulations at distances 1m, 2m and 5m show that the optimal rates are much higher (64x, 43x, and 32x, respectively) than the 54 Mbps cap applied in standard. Thus, we argue that broadcast is no longer a control-plane relic: it is time for Wi-Fi standards to treat it as a high-throughput data-plane citizen.