PATTON:为生产级LLM服务启用商用PIM
PATTON: Enabling Commodity PIM for Production LLM Serving
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
PATTON通过分层颗粒分配和提交区机制,在不修改PIM处理单元的前提下,实现商用PIM上生产级LLM服务的高效KV缓存管理,平均加速1.95倍并提升能效4.83倍。
AI中文摘要:
处理中内存(PIM)在加速内存受限的解码注意力方面前景广阔,但仅靠注意力加速不足以满足生产级LLM服务需求,因为引擎会动态分配、填充、共享、缓存和回收逻辑KV缓存块。在商用PIM上支持这一生命周期需要高效的物理内存分配、块到地址映射和命令生成。对于值缓存,这些需求在GEMV效率、单令牌写入效率和内存容量之间造成了根本性冲突:GEMV优化的布局将新生成的值向量分散到各行,导致写入成本高昂,而更细粒度的内存共享虽能提高容量利用率,却会碎片化GEMV归约。我们提出PATTON,一个将生产级LLM服务引擎与商用PIM集成的PIM运行时。PATTON引入分层颗粒分配:块大小的键和值颗粒与逻辑令牌块一一对应,固定其物理位置和命令,而更粗的颗粒将块分组以实现高效GEMV执行和内存利用。提交区(Commit Zone)暂存部分值块,以实现高效的单令牌写入,随后再将其提交到GEMV优化的位置。PATTON跟踪这些位置以生成KV缓存写入和QK转置/SV命令。在注意力执行和运行时引起的预填充重计算中,与评估基线相比,PATTON平均实现1.95倍加速和4.83倍能效提升,无需修改PIM处理单元,并保持与vLLM中原生GPU KV缓存相当的KV缓存命中率。
英文摘要:
Processing-in-Memory (PIM) is promising for accelerating memory-bound decode attention, but attention acceleration alone is insufficient for production LLM serving, where engines dynamically allocate, populate, share, cache, and reclaim logical KV cache blocks. Supporting this lifecycle on commodity PIM requires efficient physical memory allocation, block-to-address mapping, and command generation. For the Value cache, these requirements create a fundamental conflict among GEMV efficiency, single-token write efficiency, and memory capacity: GEMV-optimized layouts scatter newly generated Value vectors across rows, making writes costly, while finer-grained memory sharing improves capacity utilization but fragments GEMV reductions. We present PATTON, a PIM runtime that integrates production LLM serving engines with commodity PIM. PATTON introduces hierarchical granule allocation: block-sized Key and Value granules map one-to-one to logical token blocks, fixing their physical placements and commands, while coarser granules group blocks for efficient GEMV execution and memory utilization. A Commit Zone stages partial Value blocks for efficient single-token writes before committing them to GEMV-optimized locations. PATTON tracks these placements to generate KV cache writes and QK-transpose/SV commands. Across attention execution and runtime-induced prefill recomputation, PATTON achieves an average 1.95x speedup and 4.83x higher energy efficiency over evaluated baselines, requires no PIM processing-unit modifications, and maintains a KV cache hit rate comparable to the native GPU KV cache in vLLM.