KVBoost:用于高效大语言模型推理的、基于偏差引导重计算的分块级键值缓存复用方案
KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference
浏览论文内容
中文总结 AI 辅助
KVBoost是适用于HuggingFace兼容解码器模型的分块级KV缓存复用系统,通过双哈希键方案与两类重计算策略等,在不损失准确率的前提下,将Qwen2.5-3B的首token生成时间缩短4.49倍,性能优于前缀缓存。
中文摘要 AI 辅助
基于Transformer的大语言模型(LLM)因需为每个请求重新计算键值(KV)张量而产生较高的预填充延迟。现有的前缀缓存系统可降低该成本,但要求提示词共享连续的前导前缀,当共享内容出现在任意位置时效果受限。本文提出KVBoost,这是一种适用于HuggingFace兼容解码器模型的分块级KV缓存复用系统,支持无论内容位置如何均可复用。KVBoost引入双哈希键方案,将位置标识(前缀哈希)与内容标识(内容哈希)分离,支持精确与近似缓存匹配。针对独立缓存分块导致的注意力边界误差,KVBoost采用两种修复策略:SelectiveRecompute(选择性重计算,即重新编码边界区域)与CacheBlendRecompute(缓存混合重计算,即经探测后识别并重计算高偏差token)。该系统还集成了非对称KV量化(int8/int4)、自适应分块边界划分,以及在固定内存预算下基于重要性加权的淘汰策略。在Qwen/Qwen2.5-3B模型上,基于1000个错误定位样本的评估显示,KVBoost将首token生成时间缩短4.49倍(142.4毫秒对比639.1毫秒),性能较前缀缓存提升16%,且准确率无损失(99.2%对比99.1%)。KVBoost提供了一种实用的、受内存约束的推理加速层,兼容基于RoPE的模型且无需修改架构。
英文摘要
Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request. Existing prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, limiting effectiveness when shared content appears at arbitrary positions. We present KVBoost, a chunk-level KV cache reuse system for HuggingFace-compatible decoder models that enables reuse regardless of content position. KVBoost introduces a dual-hash keying scheme that separates positional identity (prefix hash) from content identity (content hash), supporting both exact and approximate cache matches. To address attention boundary errors from independently cached chunks, KVBoost employs two repair strategies: SelectiveRecompute, which re-encodes boundary regions, and CacheBlendRecompute, which identifies and recomputes high-deviation tokens after a probe pass. The system further incorporates asymmetric KV quantization (int8/int4), adaptive chunk boundary splitting, and importance-weighted eviction under a fixed memory budget. Evaluated on Qwen/Qwen2.5-3B over 1,000 bug-localization samples, KVBoost achieves a 4.49x reduction in time-to-first-token (142.4 ms vs.\ 639.1 ms) and outperforms prefix caching by 16%, with no loss in accuracy (99.2% vs.\ 99.1%). KVBoost provides a practical, memory-bounded inference acceleration layer compatible with RoPE-based models without architectural modification.