发表机构
Cisco Systems(思科系统公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对无线接入点的资源竞争问题,定义网络感知可部署性标准,通过基准测试揭示模型在AP与Raspberry Pi 5上的性能差异,为AP同时承载网络与ML workload提供部署依据。
AI 中文摘要
企业无线接入点(AP)是用于预测机器学习(ML)的有潜力平台,但其核心职责仍是提供无线连接与网络服务。预测推理必须与数据包处理、Wi-Fi及物联网无线电操作、客户端管理共享AP的CPU和内存。这种资源竞争带来两个风险:在代理硬件上表现良好的模型可能在目标AP上速度过慢,而单独适配的模型仍可能在负载下降低网络服务质量。我们通过两个阈值定义“网络感知可部署性”:一是模型及其执行路径在目标AP上的合格性,二是其在数据包服务和预测约束下的执行配置文件验证。我们的基准测试显示,边缘测试床无法可靠捕获目标设备行为。在匹配的构件和服务设置下,5种模型实现方案在AP上的运行速度比在Raspberry Pi 5上慢6.1至19.1倍,峰值内存使用差异最高达22%。此外,两个规模相似的预测基础模型在AP延迟上的差异达19倍。当在网络饱和条件下以30秒的周期通过13条并行流部署较小模型时,默认执行会使p99往返时间(RTT)增加76%,吞吐量降低7.06%。若要将AP同时用于网络和ML工作负载,理解这些权衡对于实时部署至关重要。
英文摘要
Enterprise wireless access points (APs) are promising platforms for predictive machine learning (ML), but their primary responsibility remains providing wireless connectivity and network services. Predictive inference must therefore share an AP's CPU and memory with packet processing, Wi-Fi and IoT radio operations, and client management. This resource contention creates two risks: a model that performs well on proxy hardware may be too slow on the target AP, while a model that fits in isolation may still degrade network services under load. We define \textit{network-aware deployability} using two gates: qualification of the model and its execution path on the target AP, followed by validation of its execution profile under packet-service and forecasting constraints. Our benchmarks show that edge testbeds do not reliably capture target behavior. Across matched artifacts and serving settings, five model implementations run 6.1--19.1$\times$ slower on an AP than on a Raspberry Pi~5, while peak memory usage differs by up to 22\%. Moreover, two forecasting foundation models of similar size differ in AP latency by 19$\times$. When serving a smaller model across 13 parallel streams at a 30~s cadence under network saturation, default execution increases p99 round-trip time (RTT) by 76\% and reduces throughput by 7.06\%. Understanding these trade-offs is essential for live deployment if we aim to use APs for both networking and ML workloads.
Comments5 pages, 2 figures, 1 table