KAP:弥合大语言模型系统中知识选择与运行时消耗的差距
KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems
浏览论文内容
中文总结 AI 辅助
研究大语言模型系统中知识选择与运行时消耗的差距,提出知识访问规划(KAP)范式,通过建立通用中间表示编译知识信号控制KV访问,以GraphSpec实例化,改变长上下文生成扩展轨迹,提升运行时消耗效率。
中文摘要 AI 辅助
现代大语言模型系统越来越依赖知识选择过程来生成高价值的结构化先验信息,如排序证据、图拓扑、多模态对齐和置信信号。然而,大语言模型服务从根本上忽略了这种丰富的结构:一旦这些信号被序列化到提示中,后端只能观察到一个扁平的令牌序列,在解码期间强制对整个键值(KV)状态进行密集且统一的消耗。我们将这种架构不匹配称为知识选择 - 运行时消耗(KSRC)差距:更丰富的上下文会扩大全提示KV占用空间和解码时的内存流量,即使推理仅依赖上下文的一小部分,也会增加延迟并降低吞吐量。为了弥合这一差距,我们提出了知识访问规划(KAP),这是一种范式转变的执行抽象,将结构化知识先验从被动的提示构建提示提升为一等的物理执行工件。KAP建立了一个通用的中间表示(IR)——运行时访问计划,它编译结构化知识信号以控制物理KV访问,而不改变逻辑提示语义、模型权重或训练过程。通过这个IR,KAP将大语言模型服务从令牌感知的上下文消耗转变为计划驱动的、知识感知的运行时消耗。我们用GraphSpec实例化KAP,它是一个编译器 - 执行器实现,将结构化知识选择连接到一个大语言模型服务后端。我们推导了计划引导执行的正加速机制的相边界模型。在4K - 128K长上下文问答工作负载中,GraphSpec保持与全上下文解码相当的答案质量,同时将物理KV消耗与提示长度解耦,在128K时将提案时的KV访问减少到源KV状态的5.5%,并从根本上改变了长上下文生成的扩展轨迹。
英文摘要
Modern LLM systems increasingly rely on knowledge-selection processes that produce high-value structured priors, such as ranked evidence, graph topology, multimodal alignment, and confidence signals. Yet LLM serving remains fundamentally oblivious to this rich structure: once such signals are serialized into a prompt, the backend observes only a flat token sequence, forcing dense and uniform consumption of the full key-value (KV) state during decoding. We term this architectural mismatch the Knowledge Selection-Runtime Consumption (KSRC) gap: richer contexts enlarge the full-prompt KV footprint and decode-time memory traffic, increasing latency and degrading throughput even when reasoning depends on only a small fraction of the context. To bridge the gap, we propose Knowledge Access Planning (KAP), a paradigm-shifting execution abstraction that elevates structured knowledge priors from passive prompt-construction hints into first-class physical execution artifacts. KAP establishes a universal intermediate representation (IR)-the runtime access plan-which compiles structured knowledge signals to govern physical KV access without altering logical prompt semantics, model weights, or training procedures. Through this IR, KAP shifts LLM serving from token-aware context consumption to plan-driven, knowledge-aware runtime consumption. We instantiate KAP with GraphSpec, a compiler-executor realization connecting structured knowledge selection to an LLM serving backend. We derive a phase-boundary model for the positive-speedup regime of plan-guided execution. Across 4K-128K long-context QA workloads, GraphSpec maintains answer quality comparable to full-context decoding while decoupling physical KV consumption from prompt length, reducing proposal-time KV access to 5.5% of source KV state at 128K, and fundamentally shifting the scaling trajectory of long-context generation.
发表机构
- QiYuanLab(启元实验室)
机构由 AI 辅助整理,请以论文原文为准。