发表机构
Florida State University(佛罗里达州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型处理长提示时的计算和内存瓶颈,提出SALT框架,将句子关键词组织成按句子频率排序的字典树,通过多锚点检索等方式保留文档主题,降低长上下文提示的预填充成本,还可与KV缓存方法结合。
AI 中文摘要
随着大语言模型处理的提示越来越长,计算和KV缓存内存成本成为推理系统的主要瓶颈。现有输入级提示压缩方法按标量相关性分数对句子排序,将文档视为单词和句子的无结构集合,在预算紧张时会导致主题崩溃。为此,我们提出了SALT,一个与模型无关的提取框架,它将每个句子的关键词组织成一个按句子频率排序的字典树,这是文档主题结构的轻量级、可重用代理。基于字典树的组织平滑了内存分配,防止主导主题垄断预算。多锚点检索可激活任何深度由查询关键词标记的字典树节点,且字典树在对话轮次中持续存在,支持多轮使用而无需重新编码文档。通过保留文档主题,SALT降低了长上下文提示的预填充计算和内存成本,同时可与针对解码时延迟和内存的KV缓存方法组合使用。
英文摘要
As large language models (LLMs) process increasingly longer prompts, computation and KV-cache memory costs have emerged as major bottlenecks in inference systems. Existing input-level prompt compression methods address this, but rank each sentence by a scalar relevance score, treating the document as an unstructured pool of words and sentences. Under tight budgets, this causes theme collapse, where the dominant theme(s) of a document consumes the budget, discarding less-frequent yet task-relevant themes. Preserving thematic coverage instead requires allocating the budget across recurring themes rather than scoring sentences in isolation. To this end, we propose SALT, a model-agnostic extractive framework that organizes per-sentence keywords into a trie ordered by sentence frequency (SF), a lightweight, reusable proxy for document thematic structure. This trie-based organization smooths memory allocation and prevents dominant themes from monopolizing the budget. Multi-anchor retrieval activates trie nodes labeled by query keywords at any depth, and the trie persists across dialogue turns, supporting multi-turn use without re-encoding the document. By preserving document themes, SALT reduces the prefill computation and memory cost of long-context prompts while remaining composable with KV-cache methods that target decoding-time latency and memory.