发表机构
The University of Hong Kong; Imperial College(香港大学; 帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对边缘LLM推理的延迟-内存权衡问题,提出混合自回归-推测式框架BALANCE,通过多项式时间算法优化资源分配,提升了任务吞吐量与服务用户数。
AI 中文摘要
边缘推理是为下一代移动网络提供大语言模型(LLM)推理服务的有前景范式。LLM推理主要依赖两种方法:自回归解码(AD)按顺序生成输出令牌,导致延迟较长;推测式解码(SD)通过使用小型语言模型(SLM)生成多个草稿令牌供LLM验证来加速推理,但会产生额外的内存开销。由于存在延迟-内存权衡,在边缘计算资源有限的情况下,两种方法单独使用均无法高效满足异构用户需求。为应对这一挑战,我们提出用于边缘LLM推理的混合自回归-推测式推理框架BALANCE。在BALANCE中,边缘服务器同时托管SLM和LLM,为每个用户分配AD或SD模式,并同时执行这两种模式。为最大化服务用户数量,我们构建任务吞吐量最大化问题,在用户延迟要求和服务器内存约束下,联合确定用户调度以及AD与SD间的计算资源分配。由于该问题为NP难问题,我们开发了一种多项式时间算法,将原问题转化为两个子问题,获得具有常数近似保证的次优解。实验表明,BALANCE始终优于传统AD和SD,并显著提升任务吞吐量。
英文摘要
Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decoding (AD) generates output tokens sequentially, resulting in long latency; Speculative decoding (SD) accelerates inference by using a small language model (SLM) to generate multiple draft tokens for LLM verification, but incurs extra memory costs. Due to this latency-memory tradeoff, neither approach alone can efficiently serve users with heterogeneous demands under limited edge computing resources. To address this challenge, we propose a hybrid autoregressive-speculative inference (BALANCE) framework for edge LLM inference. In BALANCE, an edge server hosts both an SLM and an LLM, admits users, assigns each admitted user to the AD or SD mode, and performs the two modes simultaneously. To maximize the number of served users, we formulate a task throughput maximization problem to jointly determine user admission and computing resource allocation between AD and SD under user latency requirements and server memory constraints. Since the problem is NP-hard, we develop a polynomial-time algorithm that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee. Experiments demonstrate that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.
Comments15 pages, 13 figures