发表机构
AllSpice Inc.; Boston University; Proof Trading; City College of New York; Harvard University(AllSpice公司; 波士顿大学; Proof Trading公司; 纽约城市学院; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文对上下文压缩开展形式化研究,构建含两个博弈的框架,证明上下文生成博弈与单向通信复杂性等价,还通过案例评估Anthropic的上下文压缩端点性能。
AI 中文摘要
大型语言模型(LLMs)具有受限的上下文窗口,上下文窗口是LLMs单次推理可处理的最大输入规模。AI智能体调用LLM时,依赖名为上下文压缩的过程将自身状态适配到上下文窗口内。尽管上下文压缩应用广泛,但几乎未受到形式化分析。本文启动对上下文压缩的形式化研究,首先引入由两个博弈构成的框架,该框架捕获当代AI智能体实际采用的两种上下文压缩算法策略:上下文选择博弈建模选择智能体累积状态子集进行保留的上下文压缩算法;上下文生成博弈建模以任意有限长度消息总结智能体状态的上下文压缩算法。随后证明上下文生成博弈与单向通信复杂性等价,在目标误差内回答一组查询所需的最小上下文压缩预算,等于对应通信问题在相同误差下的单向通信复杂性,因此通信复杂性的已知边界可直接迁移至上下文压缩。还证明上下文选择博弈对应受限类别的单向通信协议,选择与生成间的差距即两类通信协议间的差距,且存在一组查询使得生成所需预算严格少于选择。该等价关系还可用于衡量已部署上下文压缩算法在特定查询上相对于最优策略的性能,作为示例,本文呈现评估Anthropic上下文压缩端点在集合成员查询上的案例研究。
英文摘要
Large Language Models (LLMs) have a bounded context window. The context window is the maximum input size an LLM can consume for a single inference. AI agents rely on a process called context compaction to fit their state within the context window when calling an LLM. Despite its ubiquity, context compaction has received essentially no formal analysis. In this paper, we initiate a formal study of context compaction. We first introduce a framework consisting of two games that capture the two algorithmic strategies for context compaction used by contemporary AI agents in practice. The Context Selection Game models context compaction algorithms that select a subset of an agent's accumulated state to retain. The Context Generation Game models context compaction algorithms that summarize an agent's state by an arbitrary message of bounded length. We then prove an equivalence between the Context Generation Game and one-way communication complexity. The minimum context compaction budget for answering a set of queries within a target error is equal to the one-way communication complexity of the induced communication problem at the same error. Known bounds from communication complexity therefore transfer directly to context compaction. We also show that the Context Selection Game corresponds to a restricted class of one-way communication protocols. Any gap between selection and generation is therefore a gap between two classes of communication protocols. We prove that there exists a set of queries for which generation needs strictly less budget than selection. The equivalence between the Context Generation Game and one-way communication also lets us measure how well a deployed context compaction algorithm performs on a query relative to the optimal strategy. As an example, we present a case study that evaluates Anthropic's context compaction endpoint on set membership queries.
Comments21 pages, 2 figures, Preliminary version