发表机构
Corbenic AI(科尔贝尼克人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出galahad-kv内存层,将KV状态存于加密NVMe磁盘,在5000万token测试中,加载速度是重算的2.8-4.3倍,能耗降8.8-12.3倍,12B和31B模型对久远事实准确率达82%、98%。
AI 中文摘要
大型语言模型仅能利用其上下文窗口内的文本,且每次发送提示时都会为该提示重新计算内部键值(KV)状态。我们测试了一个名为galahad-kv的公开内存层,该层将每个约16000个token的块的KV状态保存到加密的本地NVMe磁盘中,并在之后逐字节精确地加载回来,无需重新计算。我们在5000万个token的真实公开文本上运行该层,通过vLLM在一块NVIDIA H100上部署,使用Gemma 4 12B和Gemma 4 31B模型。我们探测的每个块都从加密存储中加载,未进行任何重新计算(在0至5000万token的深度下,两个模型的100次探测均成功)。加载一个块的速度是重新计算的2.8至4.3倍,GPU能耗降低8.8至12.3倍,且在整个5000万token的数据流中,GPU内存保持稳定。当被问及数百万个token之前植入的事实时,12B模型100次中有82次给出正确答案,31B模型100次中有98次给出正确答案,两个模型均未编造答案。该方案的局限性如下:这是对存储状态的复用,而非更宽的注意力窗口,每次仅加载一个块,问题回答的效果取决于模型;写入内存是一次性成本,存储需要TB级的本地NVMe磁盘空间。我们描述了测试协议,该协议旨在抵御常见的长上下文基准测试作弊方式,并提供了单GPU复现方案,使用公开软件且该包为免费许可。
英文摘要
A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B. Every block we probed was loaded back from the encrypted store with no recompute (100 of 100, at depths from 0 to 50M tokens) on both models. Loading a block was 2.8x to 4.3x faster than recomputing it and used 8.8x to 12.3x less GPU energy, and GPU memory stayed flat over the whole 50M-token stream. Asked about facts planted millions of tokens earlier, the 12B model gave the right answer 82 times out of 100 and the 31B model 98 times out of 100. Neither model made up an answer. The limits are as follows. This is reuse of stored state, not a wider attention window: one block is loaded at a time, and how well a question is answered depends on the model. Writing the memory is a one-time cost, and the store takes terabytes of local NVMe disk. We describe the test protocol, which is built to resist common ways of gaming long-context benchmarks, and give a single-GPU reproduction that uses public software and a free licence for the package.