MANE:一种用于深度神经网络边缘卸载的多路径自适应网络
MANE: A Multi-Path Adaptive Network for Edge Onloading of Deep Neural Networks
浏览论文内容
中文总结 AI 辅助
MANE提出多路径自适应网络框架,通过多路径尾部架构、三阶段训练与滞后调度器,在边缘服务器上实现动态精度-吞吐量权衡,显著提升多设备并发下的SLO满足率与精度。
中文摘要 AI 辅助
分割计算构成了一种广泛使用的分布式推理方法,其中轻量级头部模型被卸载到设备上,而较重的尾部模型驻留在边缘服务器上,利用现代片上系统日益增长的算力,同时减轻服务器负载。随着智能办公室等智能室内环境中各类物联网设备日益增多,单个边缘服务器必须同时协助多个设备,这些设备竞争同一共享推理资源。若缺乏管理此共享负载的原则性机制,服务器将迅速过载,导致延迟服务等级协议违规,并使服务器辅助推理失效。在本工作中,我们提出MANE,一种分布式推理框架,为服务器配备多路径尾部架构,实现运行时动态的精度-吞吐量权衡。通过引入新颖的多路径模型架构、包含联合头部网络蒸馏损失的三阶段训练方案,以及具有公平设备回退策略的基于滞后的调度器,MANE在多达40个并发设备下,在最先进的卸载方法完全失败的场景中,维持超过80%的SLO满足率,同时比设备端替代方案保持高出6个百分点的精度。
英文摘要
Split computing constitutes a widely used distributed inference approach, where a lightweight head model is onloaded onto the device and a heavier tail model resides on an edge server, leveraging the growing computational capabilities of modern System-on-Chips while alleviating server load. As intelligent indoor environments such as smart offices grow increasingly populated with diverse IoT devices, a single edge server must simultaneously assist multiple devices, each competing for the same shared inference resources. Without a principled mechanism to manage this shared load, the server is quickly overwhelmed, causing latency SLO violations and rendering server-assisted inference ineffective. In this work, we present MANE, a distributed inference framework that equips the server with a multi-path tail architecture, enabling a dynamic accuracy--throughput trade-off at runtime. By introducing a novel multi-path model architecture, a three-stage training scheme featuring a Joint Head Network Distillation loss and a hysteresis-based scheduler with an equitable device-fallback policy, MANE maintains over 80% SLO satisfaction rate where state-of-the-art onloading methods fail completely, while preserving accuracy 6pp higher than on-device alternatives, across up to 40 concurrent devices.
发表机构
- National Technical University of Athens(雅典国家技术大学)
- Samsung AI Center(三星人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。