Minima-KV:采用混合格式分页注意力的保留保留KV缓存压缩技术
Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention
查看机构详情
- Minima AI, Inc.(Minima AI公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对长上下文LLM的KV缓存瓶颈,提出Minima-KV混合格式分页注意力压缩方案,在Qwen3.6-27B模型上实现3.50倍KV压缩,性能与密集控制相当,无需密集影子。
中文摘要 AI 辅助
键值(KV)缓存是长上下文大语言模型(LLM)服务中的主要容量和带宽瓶颈。本文提出Minima-KV,一种用于混合格式分页注意力的保留保留层级结构:近期且受保护的锚点页面以FP8格式保留,较旧的非锚点页面迁移至打包TQ3格式,所有活跃请求的页面均保持可寻址。特定格式的内核计算部分注意力状态,并通过全局归一化的在线Softmax合并进行组合,无需缓存大小的密集影子即可实现直接异构解码。在单块96-GB NVIDIA RTX PRO 6000 Blackwell GPU上,针对不同配置的Qwen3.6-27B模型部署,每活跃令牌的注意力KV为18.3 KiB,相较于BF16压缩率达3.50倍,相较于FP8压缩率达1.75倍。在16K RULER针插干草堆任务中,其实例化质量表现与密集控制相当;在包含503个问题的LongBench v2数据集上,16K、32K、64K上下文下的性能偏差分别为-0.80、-0.60、-0.40个百分点。针对两个59008令牌请求的单对直接解码测试显示,活跃KV压缩率达3.625倍,吞吐量为对照模型的0.9821倍,所有16个全注意力层均无回退,且未保留密集影子。这些结果为压缩长上下文状态同时不驱逐活跃请求KV页面提供了实用的混合格式路径。
英文摘要
The key-value (KV) cache is a primary capacity and bandwidth bottleneck in long-context LLM serving. We present Minima-KV, a retention-preserving hierarchy for mixed-format paged attention. Recent and protected Anchor pages remain in FP8, while older non-anchor pages move to packed TQ3; every live-request page remains addressable. Format-specific kernels compute partial attention states and combine them through a globally normalized online-softmax merge, enabling direct heterogeneous decode without a cache-sized dense shadow. Across separate, configuration-bound Qwen3.6-27B profiles on a single 96-GB NVIDIA RTX PRO 6000 Blackwell GPU, deployment accounting reports 18.3 KiB of attention KV per live token, corresponding to 3.50x compression relative to BF16 and 1.75x relative to FP8. A materializing quality profile matches its dense control on 16K RULER needle-in-a-haystack tasks. On the same 503-question LongBench v2 set, measured deltas are -0.80, -0.60, and -0.40 percentage points at 16K, 32K, and 64K. A separate single-pair direct-decode canary with two 59,008-token requests measures 3.625x active-KV compression and 0.9821x throughput relative to its control, routes all 16 full-attention layers without fallback, and retains no dense shadow. These results establish a practical mixed-format path for compressing long-context state without evicting live-request KV pages.