TopoCompress:基于图连接语义轨迹的长上下文压缩
TopoCompress: Long Context Compression via Graph-Wired Semantic Trajectories
查看机构详情
- Iowa State University(爱荷华州立大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对现有长上下文压缩方法的缺陷,本文提出无需训练、与模型无关的TopoCompress框架,通过语义片段打分与图传播实现压缩,在5项任务中表现优于基线,压缩预算仅为最强基线的1/4、时间比最快基线缩短1.41倍。
中文摘要 AI 辅助
长上下文压缩对于降低大语言模型推理的成本和延迟至关重要,但现有方法存在重要证据碎片化、需要额外训练或对齐、且往往依赖目标模型实现有效压缩的问题。本文提出TopoCompress,这是一种无需训练且与模型无关的框架,通过选择连贯的语义片段来压缩长上下文。TopoCompress首先结合稠密查询相关性、词汇查询相关性与语义加速对每个片段打分,随后构建混合图,基于语义相似性和序列邻接关系连接各片段,并在图上传播由查询引导的相关性分数。在HotpotQA、2WikiMQA、MuSiQue、Qasper和MultiFieldQA-en这5项长上下文任务中,TopoCompress始终优于强大的压缩基线;值得注意的是,它在使用仅为最强基线1/4的压缩预算时,达到了与该基线相当的性能,且相比最快基线,压缩时间缩短了1.41倍。
英文摘要
Long-context compression is essential for reducing the cost and latency of large language model inference. However, existing methods can fragment important evidence, require additional training or alignment, and often depend on the target model for effective compression. We introduce TopoCompress, a training-free and model-agnostic framework that compresses long contexts by selecting coherent semantic spans. TopoCompress first scores each span using dense and lexical query relevance together with semantic acceleration. It then constructs a hybrid graph that connects spans based on semantic similarity and sequential adjacency, and propagates the query-guided relevance scores over the graph. Across five long-context tasks-HotpotQA, 2WikiMQA, MuSiQue, Qasper, and MultiFieldQA-en-TopoCompress consistently outperforms strong compression baselines. Notably, TopoCompress achieves performance comparable to the strongest baseline while using a 4x smaller compression budget, and provides a 1.41x smaller compression time over the fastest baseline.