发表机构
Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
InferScale是原生GPU的LLM内存系统,通过可复用KV状态替代重复提示词预填充,引入Chunked RoPE与上下文窗口编码,在LoCoMo数据集上显著降低TTFT、提升吞吐量。
AI 中文摘要
大语言模型越来越多地结合持久化个性化上下文部署,例如在用户多个请求间共享的累积记忆轮廓或长对话历史。生产级内存系统(如Mem0、MemGPT和Zep)会检索该内存的相关子集并注入提示词,迫使服务引擎重复预填充相同内容。随着检索预算增加,首次生成令牌时间(TTFT)会上升,即便底层内存在请求间被复用。我们提出InferScale,这是一种原生GPU的LLM内存系统,用可复用的KV状态替代重复的提示词预填充。InferScale预计算每个记忆事实的KV表示,将其与语义嵌入一同存储在GPU上,在服务时检索相关事实,并将其KV直接注入vLLM的分页缓存。为支持旋转位置嵌入(RoPE)下动态组装的内存,我们引入Chunked RoPE,它存储旋转前的键并在注入时应用服务时的位置。不过,独立编码记忆事实会遗漏联合预填充时可用的跨事实上下文,我们通过上下文窗口编码缓解该问题:将每个记忆事实与一小段前置对话上下文一同编码,仅缓存目标事实的KV。InferScale通过vLLM的KV-connector接口实现,无需修改引擎或对模型微调。在LoCoMo数据集上针对三个开放权重模型的实验显示,当检索预算增加时,InferScale能保持TTFT近乎恒定:在k=50时,它将TTFT降低72%-79%(3.6-4.8倍),在无服务时重新计算的情况下,准确率为60.3%,而Mem0为63.3%,在并发负载下实现3.7-4.5倍的吞吐量。因此,可复用的KV状态将内存条件化服务延迟与检索上下文大小解耦,同时保持应用质量。
英文摘要
Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests. Production memory systems (e.g., Mem0, MemGPT, and Zep) retrieve a relevant subset of this memory and inject it into the prompt, forcing the serving engine to repeatedly prefill the same content. As the retrieval budget grows, time-to-first-token (TTFT) increases even though the underlying memory is reused across requests. We present InferScale, a GPU-native LLM memory system that replaces repeated prompt prefilling with reusable KV state. InferScale precomputes each memory fact's KV representation, stores it alongside a semantic embedding on the GPU, retrieves relevant facts at serving time, and injects their KV directly into vLLM's paged cache. To support dynamically assembled memories under rotary position embeddings, we introduce Chunked RoPE, which stores keys before rotation and applies their serving-time positions during injection. However, encoding memory facts independently omits the cross-fact context available during joint prefilling. We mitigate this with Context-Window Encoding, which encodes each memory fact together with a small window of preceding conversation context while caching only the target fact's KV. InferScale is implemented through vLLM's KV-connector interface, requiring neither engine modifications nor model fine-tuning. Across three open-weight models on LoCoMo, InferScale keeps TTFT nearly constant as the retrieval budget increases: at k=50 it reduces TTFT by 72-79% (3.6-4.8x), achieves 60.3% accuracy versus 63.3% for Mem0 without serving-time recomputation, and delivers 3.7-4.5x the throughput under concurrent load. Reusable KV state thus decouples memory-conditioned serving latency from retrieved-context size while preserving application quality.