arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PTStore(前缀张量存储):用于高吞吐量推理服务的分布式前缀缓存与复制

PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving

Meghana Maghyastha, Robert Underwood, Randal Burns, Bogdan Nicolae

arXiv 2607.22648首次发表:更新:

发表机构

Johns Hopkins University; Argonne National Laboraotory(约翰霍普金斯大学; 阿贡国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PTStore 受 CDN 客户端缓存设计启发,通过分发和复制流行张量形成可重用 KV 缓存前缀来加速推理,减少缓存访问延迟与负载不平衡,能大幅扩大 KV 缓存大小,使长文本问答推理效率比基线高 5 至 6 倍。

AI 中文摘要

受内容分发网络(CDN)中客户端缓存设计的启发,PTStore 对构成可重用键值(KV)缓存前缀的流行张量进行分发和复制,这是加速推理的现有方法的主要技术。这减少了访问 KV 缓存的延迟,并减轻了包含流行张量的服务器上大量请求造成的负载不平衡。此外,由于去中心化,PTStore 能将用于大语言模型(LLM)推理的 KV 缓存大小扩大几个数量级。结果表明,PTStore 对长文本问答数据集执行推理的效率比当前基线高 5 至 6 倍,当前基线不会跨不同节点和 GPU 聚合内存,因此需要重新生成 KV 缓存。

英文摘要

Inspired by the design of client caching in Content Delivery Networks (CDNs), PTStore distributes and replicates popular tensors that form reusable KV cache prefixes, which are the main technique used by state of art approaches to accelerate inferences. This reduces the latency of accessing the KV cache and alleviates load imbalance caused by a disproportionately large number of requests on servers containing popular tensors. Furthermore, thanks to decentralization, PTStore allows the expansion of the size of the KV cache for LLM inference by orders of magnitude. As a result, PTStore can execute inferences on long passage Q\&A datasets 5-6 times more efficiently than current baselines, which do not aggregate memory across different nodes and GPUs and therefore require regenerating the KV cache.

Journal refEuro-Par 2026: 32nd International European Conference on Parallel and Distributed Computing

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑