发表机构
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; Pengcheng Laboratory; University of Chinese Academy of Sciences(中国科学院深圳先进技术研究院; 鹏城实验室; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM推理服务中KV缓存的内存瓶颈问题,提出按需KV预算框架GrowPage,通过双时间尺度摘要与页级内存管理实现更优性能-吞吐量权衡。
AI 中文摘要
长输出推理使得键值(KV)缓存成为高效大语言模型(LLM)服务的关键内存瓶颈。现有KV压缩方法通常依赖预定义的每请求预算,仅调整保留哪些KV状态,而在整个解码过程中总容量固定。然而,推理工作负载存在显著的需求变化:不同请求需要不同的KV容量,且单个请求的注意力需求在生成过程中会演变。我们提出GrowPage,这是一种将KV容量视为运行时资源的按需KV预算分配框架。GrowPage维护轻量级双时间尺度查询摘要,以捕捉近期和长期的注意力行为,并利用它们的相对注意力工作集来估计需求演变。在每个容量边界,GrowPage要么压缩当前分配内的KV状态,要么在出现更广泛需求时获取额外的物理页。通过与PagedAttention的页级内存抽象集成,GrowPage保留了连续批处理和前缀缓存。在多个模型的推理基准上进行的实验表明,与现有方法相比,GrowPage实现了更优的性能-吞吐量权衡。
英文摘要
Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-request budget and adjust only which KV states are retained, leaving the total capacity fixed throughout decoding. However, reasoning workloads exhibit substantial demand variation: different requests require different KV capacities, and the attention demand of an individual request evolves during generation. We introduce \textbf{GrowPage}, an on-demand KV budgeting framework that treats KV capacity as a runtime resource. GrowPage maintains lightweight dual-timescale query summaries to capture recent and long-term attention behaviors, and uses their relative attention working sets to estimate demand evolution. At each capacity boundary, GrowPage either compresses KV states within the current allocation or acquires an additional physical page when broader demand emerges. By integrating with PagedAttention's page-level memory abstraction, GrowPage preserves continuous batching and prefix caching. Experiments on reasoning benchmarks across multiple models show that GrowPage achieves a superior performance--throughput trade-off over existing approaches.