arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04852cs.LG

KVMem:在消费级GPU上虚拟化百万令牌规模的智能体工作空间

KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

Di Chai, Leye Wang, Zeshen Su, Zhiguo Xia, Zhihang Yu

首次发表
浏览论文内容

中文总结 AI 辅助

KVMem是一种KV上下文虚拟化系统,可在消费级GPU上将智能体工作空间虚拟化至百万令牌规模,在长上下文智能体基准测试中比基于压缩的方法实现了更高的任务效用与推理效率。

中文摘要 AI 辅助

现代大语言模型(LLM)智能体在持久化工作空间中运行,其累积的历史记录可能超出GPU键值(KV)缓存容量和模型原生上下文窗口。现有系统通常将较旧的上下文压缩为摘要,或作为文本后续检索,这要么会丢失细粒度的执行证据,要么会重复预填充模型已处理过的内容。本文提出KVMem,这是一种KV上下文虚拟化系统,可将溢出的工作空间历史作为分页KV状态,在GPU内存、主机内存和NVMe之间进行存储。KVMem采用轻量级、模型原生的注意力空间索引来选择相关历史块,并实现受模型原生上下文窗口限制的、依赖于查询的执行视图。在涵盖多达100万令牌历史的长上下文智能体基准测试(包括LongMemEval、MemoryAgentBench和AgentLongBench)上进行的广泛评估显示,与作为处理上下文溢出事实上标准的基于压缩的方法相比,KVMem通常能实现更高的任务效用和推理效率。在使用Qwen3.8-27B的DeepSWE长上下文测试中,KVMem将任务成功率从仅采用压缩的上下文管理的43.8%提升至48.4%。在本地部署评估中,KVMem在配备24GB RTX 5090笔记本GPU的普通笔记本电脑上运行Qwen3.6/3.8-27B NVFP4与MTP,虚拟化了多达100万令牌的智能体工作空间——是模型原生25.6万令牌上下文窗口的四倍。在单会话设置中,KVMem生成约50令牌/秒,为本地智能体执行提供交互响应能力。更广泛地说,通过将可寻址工作空间大小与LLM的原生上下文窗口解耦,KVMem为工作空间可超出该窗口的长期运行智能体提供了一条实用路径。

英文摘要

Modern LLM agents operate in persistent workspaces whose accumulated history can exceed both GPU KV capacity and the model's native context window. Existing systems typically compact older context into summaries or retrieve it later as text, either losing fine-grained execution evidence or repeatedly prefilling content that the model has already processed. We present KVMem, a KV-context virtualization system that preserves overflowed workspace history as paged KV state across GPU memory, host memory, and NVMe. KVMem uses lightweight, model-native attention-space indexes to select relevant historical blocks and materializes a query-dependent execution view bounded by the model's native context window. Extensive evaluations on long-context agent benchmarks spanning histories up to one million tokens, including LongMemEval, MemoryAgentBench, and AgentLongBench, show that KVMem generally achieves higher task utility and greater inference efficiency than compaction-based approaches, the de facto standard for handling context overflow. In the DeepSWE long-context test with Qwen3.8-27B, KVMem improves task success from 43.8% with compaction-only context management to 48.4%. In our local-deployment evaluation, KVMem runs Qwen3.6/3.8-27B NVFP4 with MTP on an off-the-shelf laptop equipped with a 24\,GB RTX 5090 Laptop GPU, virtualizing agent workspaces of up to 1M tokens-four times the model's native 256K-token context window. In a single-session setting, KVMem generates $\sim$50 tokens/s, providing interactive responsiveness for local agent execution. More broadly, by decoupling addressable workspace size from the LLM's native context window, KVMem provides a practical path toward long-running agents whose workspaces can grow beyond that window.

发表机构

  • Shanghai University of Finance and Economics(上海财经大学)
  • Peking University(北京大学)
  • Northwestern Polytechnical University(西北工业大学)
  • Jilin University(吉林大学)

机构由 AI 辅助整理,请以论文原文为准。

↑