EvoLP:实时边缘系统中用于模型压缩的自进化延迟预测器
EvoLP: Self-Evolving Latency Predictor for Model Compression in Real-Time Edge Systems
浏览论文内容
中文总结 AI 辅助
研究针对边缘设备资源有限及测量延迟难问题,提出EvoLP框架预测模型推理延迟,该预测器可在网络压缩中自进化提升精度,实验表明其优于现有方法,融入模型压缩框架可有效指导压缩并满足延迟约束。
中文摘要 AI 辅助
边缘设备越来越多地用于在嵌入式系统上部署深度学习应用程序。许多应用程序的实时性和边缘设备的有限资源使得针对延迟的神经网络压缩成为必要。然而,在实际设备上测量延迟具有挑战性且成本高昂。因此,本文提出了一个名为EvoLP的新颖高效框架,以准确预测边缘设备上模型的推理延迟。该预测器可以在网络压缩过程中进化以实现更高的延迟预测精度。实验结果表明,EvoLP在三个边缘设备和四个模型变体上进行评估时优于先前的最先进方法。此外,当融入模型压缩框架时,它能在满足严格延迟约束的同时,有效地指导压缩过程以提高模型精度。我们在这个https URL上开源了EvoLP。
英文摘要
Edge devices are increasingly utilized for deploying deep learning applications on embedded systems. The real-time nature of many applications and the limited resources of edge devices necessitate latency-targeted neural network compression. However, measuring latency on real devices is challenging and expensive. Therefore, this letter presents a novel and efficient framework, named EvoLP, to accurately predict the inference latency of models on edge devices. This predictor can evolve to achieve higher latency prediction precision during the network compression process. Experimental results demonstrate that EvoLP outperforms previous state-of-the-art approaches by being evaluated on three edge devices and four model variants. Moreover, when incorporated into a model compression framework, it effectively guides the compression process for higher model accuracy while satisfying strict latency constraints. We open source EvoLP at https://github.com/ntuliuteam/EvoLP.
发表机构
- School of Computer Science and Engineering, Nanyang Technological University(南洋理工大学计算机科学与工程学院)
- HP-NTU Digital Manufacturing Corporate Lab, Nanyang Technological University(南洋理工大学惠普-南洋理工数字制造联合实验室)
- HP Inc.(惠普公司)
机构由 AI 辅助整理,请以论文原文为准。