FlexEE:面向卸载感知LLM推理的自推测与KV兼容早期退出方法
FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference
浏览论文内容
中文总结 AI 辅助
FlexEE通过层级监督、自推测解码和动态隐藏状态管理,在卸载场景下实现LLM早期退出,显著加速推理且精度损失极小。
中文摘要 AI 辅助
大型语言模型(LLM)推理通常受到计算和内存的双重约束,尤其是在基于卸载的部署中,自回归解码期间模型权重需要在内存层次结构之间传输。在此设置下,减少执行的层数可以降低每令牌延迟,同时避免昂贵的权重移动。基于这一观察,我们提出了FlexEE,一个面向资源受限和基于卸载的LLM推理的早期退出框架。FlexEE通过层级的退出监督实现可靠的中间层预测,通过基于Top-K局部词汇表的自推测解码实现低成本的退出决策,并通过动态隐藏状态管理实现KV缓存正确且内存感知的执行,从而使早期退出在LLM解码中变得实用。在生成任务和下游任务中,FlexEE实现了高效的早期退出,且精度下降最小,在Llama2-7B和Llama3-8B上,在0%/50%权重卸载条件下,分别实现了高达1.27倍/3.16倍和1.25倍/2.83倍的端到端加速。
英文摘要
Large language model (LLM) inference is often constrained by both computation and memory, especially in offloading-based deployments where model weights are transferred across memory hierarchies during autoregressive decoding. In this setting, reducing the number of executed layers can lower per-token latency while also avoiding costly weight movement. Motivated by this observation, we present FlexEE, an early exiting framework for resource-constrained and offloading-based LLM inference. FlexEE makes early exiting practical for LLM decoding through layer-wise exit supervision for reliable intermediate-layer prediction, self-speculative decoding over a Top-K local vocabulary for low-cost exit decisions, and dynamic hidden state management for KV-cache-correct and memory-aware execution. Across generative and downstream tasks, FlexEE enables efficient early exit with minimal accuracy degradation, delivering up to 1.27$\times$/3.16$\times$ and 1.25$\times$/2.83$\times$ end-to-end speedups on Llama2-7B and Llama3-8B under 0\%/50\% weight offloading, respectively.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。