arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

绕开计算上限:Galahad 中的字节精确内存使 LLM 阅读成为一次性成本

Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost

Sietse Schelpe

arXiv 2609.39358首次发表:更新:

发表机构

Corbenic AI(Corbenic AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM服务中重复读取文档导致计算浪费的问题,提出Galahad内存层,通过保存和加载KV状态及文档切片,使阅读成为一次性成本,显著提升召回率并降低延迟和能耗。

AI 中文摘要

Transformer 语言模型对每个 token 执行的计算量是有界的,Infosys 前首席执行官 Vishal Sikka 最近的工作(arXiv:2507.07505)认为,这一界限限制了模型能够执行或验证的任务。我们探究在该上限之下,有多少计算预算被花费在模型已经完成的工作上。服务在请求之间是无状态的:一个模型回答关于某文档的第二个问题时,会从第一个 token 开始重新计算该文档的注意力状态。在七个真实世界数据集上,98.7% 的提示 token 是模型已经阅读过的文本。我们提出了 Galahad,一个用于 vLLM、SGLang 和此 http URL 的内存层,它使这种阅读成为一次性成本。Taliesin 保存模型对一段文本的键值(KV)状态,并在下一个包含相同字节的请求时加载它,而不是重新计算。Blaise 保留文档本身,只将问题所需的章节传递给模型。在一个将 100 个事实隐藏在 97,000 token 语料库(Gemma 4 31B)中的召回测试中,仅使用 Taliesin 就让模型能够关注整个语料库,并在 3.0 秒和每个问题 572 焦耳的情况下回答了 100 个问题中的 98 个,而同一模型在没有 Galahad 的情况下只能保留最后 12,000 个 token,回答了 100 个中的 10 个,耗时 9.3 秒,能耗 2,754 焦耳。加入 Blaise 后,模型每个问题阅读约 668 个 token,并在三种运行时上以 0.59-0.64 秒和 200-213 焦耳回答了 100 个中的 100 个;一个经过调优的 RAGFlow 流水线回答了 77 个。存储语料库是一次性成本,约 100 秒和 28 千焦,其能量在 13 个问题后即可回收。恢复的状态是位级相同的:在重启、再水化和热加载后,所有 262,144 个输出 logits 均匹配。Galahad 与我们测试的 vLLM 下所有 30 个模型都能正常工作,并且它失败时关闭:任何未通过其检查的加载都会被重新计算。这些结果共同将 LLM 服务从无状态推理转变为有状态推理。

英文摘要

A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out or verify (arXiv:2507.07505). We ask how much of the budget beneath that ceiling is spent on work the model has already done. Serving is stateless across requests: a model that answers a second question about a document recomputes the document's attention state from the first token. On seven real-world datasets, 98.7% of prompt tokens were text the model had already read. We present Galahad, a memory layer for vLLM, SGLang and llama.cpp that makes this reading a one-time cost. Taliesin saves the model's key-value (KV) state for a block of text and loads it on the next request that contains the same bytes, instead of recomputing it. Blaise keeps the documents themselves and passes the model only the section a question needs. On a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone let the model attend to the whole corpus and answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question, against 10 of 100, 9.3 s and 2,754 J for the same model without Galahad, which could hold only the last 12,000 tokens. With Blaise added, the model read about 668 tokens per question and answered 100 of 100 on all three runtimes at 0.59-0.64 s and 200-213 J; a tuned RAGFlow pipeline answered 77. Storing the corpus is a one-time cost of about 100 s and 28 kJ, whose energy is recovered after 13 questions. Restored state is bit-identical: all 262,144 output logits matched after restart, rehydration and hot-load. Galahad worked with all 30 models we tested under vLLM, and it fails closed: any load that does not pass its checks is recomputed. Together these results move LLM serving from stateless to stateful inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑