arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13153cs.ARcs.DC

面向移动边缘设备上小语言模型推理的DVFS

DVFS for Small Language Model Inference on Mobile Edge Devices

Jiesong Chen, Lixiang Han, Jiani Cao, Zhaoxi Yue, Zhenjiang Li

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出DVFSLM,一种针对移动边缘设备上小语言模型推理的DVFS设计,通过工作负载感知估计器和运行时调控器,在满足延迟截止时间的同时降低每令牌能耗,实验显示能效和QoS显著提升。

中文摘要 AI 辅助

本文提出了DVFSLM,一种用于移动边缘设备上小语言模型(SLMs)能效推理的新型动态电压频率调整(DVFS)设计。对语言模型本地执行的需求日益增长,推动了SLMs的采用,这些模型在计算可行性与良好推理性能之间取得了平衡。然而,能效仍然是一个关键挑战,因为即使是小型化的SLMs也会产生显著的能耗,影响应用质量、设备可靠性和环境可持续性。现有的针对云端大型模型或通用移动工作负载设计的DVFS解决方案,未能解决SLMs独特的工作负载特征,导致能量浪费或过度延迟。与先前工作不同,DVFSLM明确解决了两个关键挑战:1)自回归令牌生成过程中处理器频率、功耗和延迟之间的复杂相互依赖关系;2)硬件不透明性,即在协作执行期间,不同处理器(GPU、CPU和EMC)各自的功耗和延迟贡献被掩盖。为解决这些问题,DVFSLM引入了工作负载感知的功耗和延迟估计器,分析核心矩阵运算并将其与硬件元数据关联,从而能够精确估计频率调整对功耗和延迟的影响。这些估计驱动一个运行时DVFS调控器,该调控器协调GPU和EMC频率与基于配置文件的CPU频率阈值,在满足可配置的令牌生成截止时间的同时,最小化每令牌能耗。在丰富多样的SLMs上进行的大量实验表明,与最新的内置调控器相比,DVFSLM将每令牌能耗降低了高达12.4%,与最先进的GearDVFS相比降低了高达8.4%,同时分别将延迟服务质量(QoS)提高了高达93.12%和69.14%。

英文摘要

This paper presents DVFSLM, a new dynamic voltage and frequency scaling (DVFS) design for energy-efficient inference of small language models (SLMs) on mobile edge devices. The growing demand for local execution of language models has driven the adoption of SLMs, which balance computational feasibility with good inference performance. However, energy efficiency remains a critical challenge, since even miniaturized SLMs impose significant energy consumption, impacting application quality, device reliability, and environmental sustainability. Existing DVFS solutions, designed for cloud-based large models or generic mobile workloads, fail to address the unique workload characteristics of SLMs, resulting in wasted energy or excessive latency. Unlike prior work, DVFSLM explicitly addresses two key challenges: 1) the complex interdependencies of processor frequencies, power and latency across autoregressive token generations, and 2) hardware opacity, where the individual power and latency contributions from different processors (GPU, CPU and EMC) are obscured during collaborative execution. To address these, DVFSLM introduces workload-aware power and latency estimators that analyze core matrix operations and correlate them with hardware metadata, enabling precise estimations of how frequency adjustments impact power and latency. These estimations drive a runtime DVFS governor that coordinates the GPU and EMC frequencies with a profiled CPU-frequency threshold, minimizing the energy per token while satisfying configurable token-generation deadlines. Extensive experiments on a rich set of SLMs show that DVFSLM reduces the energy per token by up to 12.4% over the latest built-in governors and up to 8.4% over the state-of-the-art GearDVFS, while improving the latency quality of service (QoS) by up to 93.12% and 69.14%, respectively.

发表机构

  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

↑