AI 中文总结
该研究提出面向KV缓存的互联网愿景,主张跨云与数据中心解耦计算与KV缓存存储,将KV缓存管理作为内容分发系统,以最小化大语言模型推理的延迟与成本。
AI 中文摘要
大语言模型推理已成为覆盖智能体、检索、工具使用、代码执行及多模态推理的全球规模异构工作负载。这些工作负载天然支持来自重叠输入的上下文复用,为存储和复用上下文的KV缓存而非重新计算创造了重大机遇。然而,缩减KV缓存的模型端进展与降低计算、存储及传输成本的系统端进展,在传统云边界内独立演进。我们认为,未来推理基础设施应允许跨云与数据中心解耦计算与KV缓存存储,网络成为主动分发渠道,带宽、延迟及定价直接决定KV缓存的管理方式。我们提出面向KV缓存的互联网愿景,将KV缓存管理作为内容分发系统,在此视角下,KV缓存存储与重计算决策由模型、基础设施及应用指标驱动,以实现自适应、内容驱动的决策,从而最小化延迟与成本。
英文摘要
LLM inference has become a global-scale, heterogeneous workload spanning agents, retrieval, tool-use, code execution and multi-modal reasoning. These workloads naturally enable context reuse from overlapping inputs, creating a major opportunity to store and reuse the contexts' KV Caches instead of recomputing them. However, model-side advances that shrink the KV Cache and system-side advances that reduce compute, storage, and transfer costs are evolve independently within legacy cloud boundaries. We argue that future inference infrastructure should allow decoupling of compute and KV Cache storage across cloud and datacenters. The network becomes an active distribution channel; bandwidth, latency and pricing directly determines how the KV Cache should be managed. We propose a vision for an Internet for the KV Cache, with KV Cache management working as a content-distribution system. In this view, KV Cache storage and recompute decisions are driven by model, infrastructure, and application metrics, to enable adaptive, content-driven decisions for minimizing latency and cost.