发表机构
Kennesaw State University; Samsung Semiconductor, Inc.(肯尼索州立大学; 三星半导体公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型部署时模型超出内存的场景,NeuroPrefetcher利用MLP活动的时间局部性,通过GPU驻留预测器进行增量预取,实现了7.9-12.0倍的推理加速。
AI 中文摘要
将大语言模型部署在边缘设备上,日益受到模型大小与可用内存之间不断扩大的差距的限制。现有方法如量化、使用更小的模型以及模型卸载,能够提高有效内存限制,但这些方法仍假设模型可以被压缩或分区以适应一定的预算。本研究针对更困难的“模型超出内存”场景,即模型在整个执行过程中仍然大于常驻内存,存储成为关键路径上权重的主动来源。我们观察到,自回归解码期间的多层感知机(MLP)活动具有很强的时间局部性:约82%至85%的活跃神经元在相邻 token 之间保持不变。这意味着当前 token 所需的大多数稀疏权重已经驻留,仅需从存储中获取新需要的行。我们提出了NeuroPrefetcher,这是一个基于存储的大语言模型推理系统,通过预测性增量预取技术利用这一特性。在第0层之后,一个仅占用基础模型参数2.86%、驻留在GPU中的单一预测器,通过一次前向传播预测所有后续MLP层的稀疏活动。运行时将这些预测与常驻GPU缓冲区进行比较,并仅针对传入的增量行发起应用程序调度的NVMe读取,以显式的、模型感知的权重移动替代反应式的操作系统请求分页。在真实的统一内存边缘硬件上,NeuroPrefetcher在受限内存预算下实现了比基线方法7.9至12.0倍的加速。
英文摘要
Deploying large language models on edge devices is increasingly limited by a widening gap between model size and available memory. Existing approaches such as quantization, smaller models, and offloading can raise the effective memory limit, but they still assume that the model can be compressed or partitioned to fit within some budget. We target the harder model-exceeds-memory setting, in which the model remains larger than resident memory throughout execution and storage becomes an active source of weights on the critical path. We observe that MLP activity during autoregressive decoding has strong temporal locality: approximately 82-85% of active neurons persist from one token to the next. This means that most sparse weights needed for the current token are already resident, and only the newly needed rows must be fetched from storage. We present NeuroPrefetcher, a storage-backed LLM inference system that exploits this property through predictive delta prefetching. After layer 0, a single GPU-resident predictor, occupying 2.86% of base model parameters, predicts sparse activity for all downstream MLP layers in one forward pass. The runtime compares these predictions against resident GPU buffers and issues application-scheduled NVMe reads only for incoming delta rows, replacing reactive operating-system demand paging with explicit, model-aware weight movement. On real unified-memory edge hardware, NeuroPrefetcher achieves 7.9-12.0x speedup over llama.cpp across constrained memory budgets.
Comments10 pages, 10 figures. Accepted at the 55th International Conference on Parallel Processing (ACM ICPP 2026), Singapore. Code, training pipeline, and measured data: https://github.com/nobeldhar/NeuroPrefetcher