发表机构
Kyungpook National University(庆北国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对边缘设备推测解码中内存受限导致自适应方法吞吐量提升不足的问题,提出MemSpec内存感知运行时,通过解耦草稿选择与执行等设计,使Jetson Orin Nano上的稳态生成吞吐量平均提升40.7%
AI 中文摘要
推测解码通过使用轻量级草稿模型推测多个token,减少昂贵的目标模型解码步骤,从而加速自回归大语言模型(LLM)的推理,其有效性在很大程度上取决于草稿选择,这促使人们开发了利用输入和生成阶段差异的自适应方法。然而,在内存受限的边缘设备上,这些方法往往因草稿模型之间切换的开销而无法提高端到端吞吐量。我们发现该场景下的一个关键局限:在紧张的内存预算下,草稿选择与草稿可用性之间存在不匹配。为应对这一挑战,我们提出了MemSpec,这是一种用于边缘设备自适应推测解码的、由预测引导的内存感知运行时。MemSpec通过主动驻留工作集管理将草稿选择与执行解耦,轻量级预测器根据提示和生成上下文估计草稿有效性,而内存感知调度器则减少反应式模型加载开销。在Jetson Orin Nano上进行的实验表明,与最先进的基于多臂老虎机的自适应方法相比,MemSpec平均将稳态生成吞吐量提高了40.7%,同时接近最优上限。
英文摘要
Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to speculate multiple tokens, reducing expensive target model decoding steps. Its effectiveness depends heavily on draft selection, motivating adaptive methods that exploit variation across inputs and generation stages. On memory-constrained edge devices, however, these methods often fail to improve end-to-end throughput due to the overhead of switching between draft models. We identify a key limitation in this setting: the mismatch between draft selection and draft availability under tight memory budgets. To address this challenge, we present MemSpec, a prediction-guided, memory-aware runtime for adaptive speculative decoding on edge devices. MemSpec decouples draft selection from execution through proactive resident working-set management. A lightweight predictor estimates draft effectiveness from prompt and generation context, while a memory-aware scheduler reduces reactive model loading overhead. Experiments on a Jetson Orin Nano show that MemSpec improves steady-state generation throughput by 40.7% on average over state-of-the-art bandit-based adaptive methods while closely approaching the oracle upper bound.
CommentsPublished in LCTES 2026
Journal refProc. ACM LCTES 2026, 180-192 (2026)