发表机构
National University of Defense Technology; Peking University(国防科技大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
vToken是一种轻量级令牌级虚拟化层,可解耦逻辑令牌活性与物理块放置,在保留PagedAttention兼容性的同时,提升KV缓存回收效率与服务吞吐量,降低策略集成代码量。
AI 中文摘要
大语言模型服务面临关键内存瓶颈:KV缓存随序列长度和批量大小增长。PagedAttention使用固定大小内存块减少分配器级碎片,但近期KV驱逐算法以比块级管理更细的令牌粒度运行,这种不匹配导致块内碎片,使大量已分配KV内存无法回收。我们提出vToken,一种轻量级令牌级虚拟化层,将逻辑令牌活性与物理块放置解耦,通过令牌表间接维护稳定的逻辑令牌视图,并通过异步重打包存活令牌实现物理回收。该设计保留PagedAttention内核和CUDA Graph兼容性。我们在vLLM中实现vToken,结合H2O、Random和Scissorhands策略在多个模型上评估:与配对的Naive-Evict基线相比,vToken将每个请求保留的KV块减少27.2%至72.3%,受SLA约束的吞吐量最高提升1.37倍;在受限的活动KV预算下,它将最大可行并发数扩展最高2倍,同时将每个策略的集成占用空间从500多行代码减少到50行以下。
英文摘要
Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragmentation, but recent KV eviction algorithms operate at a token granularity finer than block-level management. This mismatch causes intra-block fragmentation, leaving a large fraction of allocated KV memory unreclaimable. We present vToken, a lightweight token-level virtualization layer that decouples logical token liveness from physical block placement. vToken maintains a stable logical token view through token-table indirection and realizes physical reclamation by repacking live tokens asynchronously. The design preserves PagedAttention kernels and CUDA Graph compatibility. We implement vToken in vLLM and evaluate it with H2O, Random, and Scissorhands across models. Compared with a paired Naive-Evict baseline, vToken reduces retained KV blocks per request by 27.2\%--72.3\% and improves SLA-constrained throughput by up to 1.37$\times$. Under a constrained active-KV budget, it extends the maximum feasible concurrency by up to 2$\times$, while reducing the per-policy integration footprint from 500+ lines to under 50.