arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16161cs.LG

LLM推理:闪速进行!

LLM Inference in a Flash!

发表机构加州大学伯克利分校 · 国际计算机科学研究所 · 劳伦斯伯克利国家实验室
查看机构详情
  • UC Berkeley(加州大学伯克利分校)
  • ICSI(国际计算机科学研究所)
  • LBNL(劳伦斯伯克利国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

Sebastian Zhao, Minseo Kim, Coleman Hooper, Luca Manolache, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM推理中内存带宽和写入耐久性挑战,提出纯整数量化与字典式KV缓存压缩算法,在Flash存内计算设备上实现高效推理,减少15倍动态KV缓存流量。

中文摘要 AI 辅助

大型语言模型(LLMs)在各类自然语言处理任务中展现出令人印象深刻的能力,而LLM推理已成为支持下游应用的关键工作负载。随着请求转向更长的序列和更重的推理负载(由检索增强生成、推理时计算扩展和长上下文应用驱动),服务LLM推理的需求正变得越来越具有挑战性。此外,这些挑战因硬件趋势而加剧,因为内存容量和通信带宽的扩展速度不及工作负载复杂性的增长。Compute-in-Flash是一种有前景的解决方案,通过将计算移至靠近内存的位置来解决内存带宽限制,并利用SSD技术的大容量优势。然而,在这些系统上部署LLM颇具挑战性,因为它们缺乏对高精度浮点运算的支持,且写入耐久性有限。在我们的工作中,我们旨在通过设计推理算法来解决这些挑战,从而在Flash存内计算设备上实现LLM推理。我们提出了一种端到端的纯整数量化方法,以消除昂贵的浮点计算。为解决有限的写入耐久性,我们设计了一种基于稀疏字典编码的字典式KV缓存压缩策略,将每个KV向量表示为静态字典向量的线性组合。这些算法改进使我们能够利用Compute-in-Flash对模型权重和KV缓存的双重优势,并最小化昂贵的数据传输操作。在Llama-3.1-8B和Qwen-2.5-7B上,我们的组合方法在将动态KV缓存流量减少15倍的同时,仅表现出有限的精度下降。

英文摘要

Large Language Models (LLMs) have shown impressive capabilities across a range of natural language processing tasks, and LLM inference has emerged as a critical workload for enabling downstream applications. The demands of serving LLM inference are becoming increasingly challenging as requests shift toward longer sequences and heavier inference, driven by retrieval-augmented generation, inference-time compute scaling, and long-context applications. Additionally, these challenges are compounded by hardware trends, as memory capacity and communication bandwidth are not scaling as fast as increases in workload complexity. Compute-in-Flash is a promising solution to address memory bandwidth limitations by moving computation close to memory, and to exploit the large capacity of SSD technologies. However, it is challenging to deploy LLMs on these systems as they lack support for high-precision floating point operations and have limited write endurance. In our work, we aim to address these challenges by designing inference algorithms to enable LLM inference on Flash compute-in-memory devices. We present an end-to-end integer-only quantization approach to eliminate expensive floating-point computations. To address the limited write endurance, we design a dictionary-based KV cache compression strategy based on sparse dictionary coding that represents each KV vector as a linear combination of static dictionary vectors. These algorithmic improvements enable us to exploit the benefits of Compute-in-Flash for both model weights and KV cache, and to minimize expensive data transfer operations. Across Llama-3.1-8B and Qwen-2.5-7B, our combined method exhibits limited accuracy degradation while reducing dynamic KV cache traffic by 15$\times$.

↑