分形KV缓存存档:用于长上下文语言模型推理的带原地检索的无损符号存储
Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference
浏览论文内容
中文总结 AI 辅助
研究长上下文语言模型推理中KV缓存的存储问题,提出用分形KV缓存存档方法,该方法无损、线性时间运行且支持随机访问与追加,经实验验证能大幅减少存档缓存,还兼具搜索索引功能,发布代码可重现结果。
中文摘要 AI 辅助
键值(KV)缓存主导着长上下文自回归推理的内存成本,越来越多的工作通过量化、逐出或卸载来压缩它。我们研究一个补充问题:一旦一个位置的KV状态被量化为码本索引,如何存储由此产生的符号流,存储层能否做更多的事情?我们重新审视了一类将符号序列串行化为低维实向量序列的压缩迭代映射码,结果表明它们形成了一种自然的存档格式,具有以下特点。该方法提供了不断增长的缓存所需的访问模式。它是无损的,运行时间为线性,支持O(1)随机访问和O(1)分摊追加。我们在具有1024个令牌上下文的GPT-2上对为该存档提供数据的量化器进行了对照研究。保持一个小的精确窗口(4个注意力汇聚点+32个最近的令牌)并对其余部分进行存档,相对于fp16缓存,每个头的残差向量量化将存档缓存减少了36 - 54倍,困惑度成本为11 - 15%,并且我们量化了明显的键/值不对称性——量化键的损害大约是量化值的4倍,这与之前低比特KV工作一致——并将其用于混合方案中的比特分配。最后,我们表明该存档同时也是一个搜索索引:近似子串查询直接在存储的向量上执行,匹配的上下文从匹配的向量中解码,而无需实际化周围的文本。我们发布了所有代码;每个数字都可以在笔记本电脑CPU上通过单个命令重现。
英文摘要
The key-value (KV) cache dominates the memory cost of long-context autoregressive inference, and a growing body of work compresses it through quantization, eviction, or offloading. We study a complementary question: once a position's KV state has been quantized to codebook indices, how should the resulting symbol stream be stored, and can the storage layer do more than store? A family of contractive iterated-map codes that serialize a symbol sequence into a sequence of low-dimensional real vectors is revisited, and it is shown that they form a natural archive format for a quantized KV cache with the following features. The method provides exactly the access pattern a growing cache requires. It is lossless, it runs in linear time, and supports O(1) random access and O(1) amortized append. A controlled study of the quantizer feeding this archive is conducted on GPT-2 with 1024-token contexts. Keeping a small exact window (4 attention sinks plus 32 recent tokens) and archiving the rest, per-head residual vector quantization reduces the archived cache by 36-54x relative to an fp16 cache at a perplexity cost of 11-15%, and we quantify a sharp key/value asymmetry - quantizing keys is roughly 4x more damaging than quantizing values, consistent with prior low-bit KV work - and use it to allocate bits in a hybrid scheme. Finally, we show the archive is simultaneously a search index: approximate substring queries execute directly on the stored vectors, and matched context is decoded from the matched vector without ever materializing the surrounding text. We further characterize the archive's operating range: truncating stored points yields a lossy regime whose distortion we localize, and probability-weighting the maps recovers arithmetic coding, exposing a trade-off between rate efficiency, random access, and memory. We release all code; every number reproduces on a laptop CPU.
发表机构
- Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。