arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向高能效DNN推理的存储与计算频率联合优化

Joint Optimization of Memory and Computing Frequency for Energy-Efficient DNN Inference

Yunchu Han, Zhaojun Nan, Sheng Zhou, Zhisheng Niu

arXiv 2608.13863首次发表:更新:

发表机构

Chongqing Polytechnic University of Electronic Technology(重庆电子工程职业学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对移动设备DNN推理的高能耗问题,联合优化存储与计算频率,提出近最优闭式解、最优传输功率解及低复杂度启发式算法,可降低设备能耗最高10.4%。

AI 中文摘要

移动设备上的深度神经网络(DNN)推理常因计算与存储资源有限而导致高延迟与高能耗。为实现高能效DNN推理,现有多数研究聚焦于动态电压频率缩放(DVFS)以调整计算频率,却极大忽略了存储频率对推理性能的影响。本文考虑存储频率与计算频率对DNN推理时间的影响,将这两种频率与通信资源联合优化,以实现高能效DNN推理。基于实际推理时间模型,我们构建优化问题,在截止期限约束下最小化所有移动设备的能耗。对于本地推理,我们通过凸优化推导得到近最优闭式解;对于给定带宽的边缘推理,我们得到传输功率的最优闭式解。此外,我们提出一种低复杂度启发式算法,可在多项式时间复杂度下有效求解整体问题。基于实测数据的仿真结果表明,所提出的本地推理近最优解在严格截止期限约束下可达到最优性能,与最优解的性能差距最高为2.5%;同时,与其他方法相比,本文提出的算法可将设备能耗显著降低最高达10.4%。

英文摘要

Deep neural network (DNN) inference on mobile devices often incurs high latency and energy consumption due to limited computing and memory resources. To enable energy-efficient DNN inference, most existing studies focus on dynamic voltage and frequency scaling (DVFS) for adjusting the computing frequency, while the impact of memory frequency on the inference performance has been greatly overlooked. In this paper, we consider the impact of memory frequency and computing frequency on DNN inference time, and jointly optimize these two frequencies together with communication resources for energy-efficient DNN inference. Based on a realistic inference time model, we formulate an optimization problem to minimize the energy consumption of all mobile devices under the deadline constraint. For local inference, we derive a near-optimal closed-form solution via convex optimization, while an optimal closed-form solution for transmission power is obtained for edge inference with the given bandwidth. Furthermore, we propose a low-complexity heuristic algorithm to effectively solve the overall problem with polynomial time complexity. Simulation results based on measured data show that the proposed near-optimal solution for local inference can achieve optimal performance under strict deadline constraints, with a performance gap of up to 2.5% compared with the optimal solution. Meanwhile, our proposed algorithm significantly reduces the energy consumption of devices by up to 10.4% compared to other methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑