发表机构
George Mason University; Global Technology Applied Research, JPMorganChase; University of Kansas(乔治梅森大学; 摩根大通全球技术应用研究部; 堪萨斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LeanStream是一款端侧LLM推理的流式推测-优化框架,通过优化计算、加载及缓存优先级实现GPU执行与存储I/O重叠,可降低内存使用并提升token生成吞吐量。
AI 中文摘要
端侧LLM推理因隐私性和响应性颇具吸引力,但在移动设备与嵌入式设备上仍具挑战性,因为模型权重远超可用DRAM。现有系统利用激活稀疏性并将权重卸载至SSD或闪存存储,却面临根本性的系统权衡:准确的稀疏执行决策需要最新上下文,而高效的计算-I/O重叠则需要早期预测。因此,现有设计要么串行执行,要么产生冗余权重获取、额外计算及大缓存开销。本文提出LeanStream,一种用于高效端侧LLM推理的流式推测-优化框架。LeanStream利用部分GPU结果逐步优化计算、加载及缓存保留优先级,实现GPU执行与存储I/O的细粒度重叠。我们在移动和嵌入式平台上实现了LeanStream。与现有端侧LLM推理系统相比,在现有工作达到的最佳吞吐量下,LeanStream将内存使用量降低4.8倍至7.5倍,同时将token生成吞吐量进一步提升1.6倍至2.1倍。
英文摘要
On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior systems exploit activation sparsity and offload weights to SSD or flash storage, but face a fundamental systems trade-off: accurate sparse execution decisions require the latest context, whereas efficient computation-I/O overlap requires early prediction. As a result, existing designs either serialize execution or incur redundant weight fetches, extra computation, and large cache overheads. We present LeanStream, a streaming speculate-and-refine framework for efficient on-device LLM inference. LeanStream progressively refines computation, loading, and cache-retention priorities using partial GPU results, enabling fine-grained overlap between GPU execution and storage I/O. We implement LeanStream on both mobile and embedded platforms. Compared with prior on-device LLM inference systems, LeanStream reduces memory usage by 4.8$\times$ to 7.5$\times$ at the best throughput achieved by prior work, while further improving token generation throughput by 1.6$\times$ to 2.1$\times$.
Comments6 pages, 13 figures. To appear in the Proceedings of the 32nd Annual International Conference on Mobile Computing and Networking (MobiCom '26)