发表机构
University of Bari Aldo Moro(巴里阿尔多·莫罗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Geodesia-KV无训练KV缓存策略,通过单调等势分配等技术实现高效KV缓存压缩,在多数据集上优于基线方法,可在单GPU上生成1M token上下文并节省大量VRAM。
AI 中文摘要
当前对KV缓存压缩性能的评估混淆了驻留位与读取带宽,且受分块教师强制的伪影影响。本文提出Geodesia-KV,一类基于单调分块精度分配、精确率失真残差及查询稀疏读取的无训练KV缓存策略,可实现适配硬件的压缩。通过合理分离驻留位与读取位并进行因果评估,Geodesia-KV显著优于其他方法:在WikiText-2的16k上下文场景下,Geodesia-KV的5比特工作点在更低比特率下实现比Qwen上的KIVI-4更低的困惑度;其compressed-Quest版本在PG-19上相比基线稀疏方法,困惑度提升且驻留位(9.83 vs 16.25比特/值)与读取率(1.95 vs 2.32比特/值)均降低。由于Geodesia-KV作为vLLM的原生GeodesiaKVCacheManager插件实现,可通过单调比特降级完全消除对密集缓存驻留的需求。经全消费硬件评估,Geodesia-KV可在单块16 GiB GPU上实现1M token的上下文生成,在Qwen、Llama、DeepSeek等主流架构上实现最高71.7%的峰值VRAM节省。
英文摘要
Current assessment of KV-cache compression performance confuses resident bits with read bandwidth and is affected by the artifacts of chunked teacher-forcing. We present Geodesia-KV, a family of training-free KV cache policies based on monotonic block-wise precision allocation, exact rate-distortion residuals, and query-sparse reading, enabling proper hardware-ready compression. With proper separation of resident and read bits and causal evaluation, we show that Geodesia-KV significantly outperforms other approaches. Specifically, on WikiText-2 with 16k context, the 5-bit operating point of Geodesia-KV results in lower perplexity at lower bitrate than KIVI-4 on Qwen. In addition, our compressed-Quest version delivers improved perplexity and reduces resident (9.83 vs 16.25 bits/value) and read rates (1.95 vs 2.32 bits/value) over baseline sparse methods on PG-19. As Geodesia-KV is implemented as native GeodesiaKVCacheManager plug-in of vLLM, Geodesia-KV fully removes the need for dense cache residency via monotonic bit demotion. With the full consumer hardware evaluation, Geodesia-KV leads to 1M-token context generation on a single 16 GiB GPU with up to 71.7% peak VRAM savings on all leading architectures (Qwen, Llama, DeepSeek).
Comments11 pages, 3 figures, 8 tables. Code and vLLM integration available at https://github.com/geodesia-ai/geodesia-kv