arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PELM:基于推测解码和动态电压频率调整的节能设备端大语言模型推理

PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling

Weisi Yang, Stephen Xia

arXiv 2609.09662首次发表:更新:

发表机构

Northwestern University(西北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对移动端 LLM 推理的功耗和热量问题,提出 PELM,结合 DVFS、推测解码和可变验证深度,实现高达 23.1% 加速和 52.4% 能耗降低。

AI 中文摘要

在边缘端移动平台上直接部署大型语言模型(LLM)正受到越来越多的关注,因为它带来了诸多好处,例如增强隐私性、个性化和降低延迟。然而,LLM 具有高计算需求,这对于资源受限的移动和边缘平台来说难以满足。除了有限的计算资源外,移动和边缘系统通常具有紧凑的外形,并且缺乏物理机制(如风扇)来散发高处理器使用率产生的热量,以防止降频和计算能力下降,而 LLM 很容易导致这种情况。为了缓解这些影响,先前的工作提出了各种功耗管理策略,例如动态电压和频率调整(DVFS),用于减少移动平台上高计算任务产生的功耗和热量。最近,针对移动 LLM 定制的 DVFS 方法也被提出。然而,这些方法主要侧重于优化硬件参数和处理器频率,在某些热受限场景下表现不足。借鉴机器学习的最新进展,我们发现并利用了关键洞察:并非所有 token 都需要全深度推理来保持高质量生成。受此启发,我们提出了 PELM,一种通过两个额外的负载特定控制旋钮来增强传统 DVFS 处理器频率调整的解决方案:1) 推测解码和 2) 可变验证深度,将优化空间扩展到多个维度,以实现更节能的设备端 LLM 推理。在跨硬件平台和数据集的广泛评估中,PELM 与最先进的功耗管理方法相比表现出优越的性能,实现了高达 23.1% 的加速和 52.4% 的能耗降低,同时保持了相当的任务性能。源代码可在以下 https URL 获取。

英文摘要

Deploying Large Language Models (LLMs) directly on mobile platforms at the edge is gaining traction due to a myriad of benefits, such as increased privacy, personalization, and reduced latency. However, LLMs have heavy computational requirements, which are difficult for resource-constrained mobile and edge platforms to fulfill. In addition to limited compute resources, mobile and edge systems often have a compact form factor and lack physical mechanisms to dissipate heat generated from high processor usage rates (e.g., fans) to prevent throttling and reduced processing power, which LLMs can easily cause. To mitigate these effects, prior works have proposed various power governing strategies, such as dynamic voltage and frequency scaling (DVFS), for reducing power and heat generation for heavy computational tasks on mobile platforms. Recently, DVFS methods tailored for mobile LLMs have also been proposed. However, these methods mostly focus on optimizing hardware parameters and processor frequencies, and they fall short under some thermally constrained scenarios. Drawing from recent advances in machine learning, we identify and take advantage of the key insight that not all tokens require full-depth inference to maintain high-quality generation. Motivated by this, we present PELM, a solution that augments traditional DVFS processor frequency tuning with two additional workload-specific knobs: 1) speculative decoding and 2) variable verification depth to expand the optimization space to multiple dimensions for more power efficient on-device LLM inference. In extensive evaluations across hardware platforms and datasets, PELM demonstrates superior performance compared to state-of-the-art power governing methods, with up to 23.1% speedup and 52.4% reduction in energy consumption, while maintaining comparable task performance. The source code is available at https://github.com/imec-nu/PELM.

CommentsAccepted to ACM/IEEE SenSys'26

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑