arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为大语言模型智能体策划始终加载的上下文:带删失反馈的容量约束 assortment 模型

Curating Always-Loaded Context for LLM Agents: A Capacitated Assortment Model with Censored Feedback

Zexuan Liu, Yuning Yang, Tiancheng Zhao

arXiv 2610.11007首次发表:更新:

发表机构

University of Illinois at Urbana-Champaign; Georgia Institute of Technology; Saint Louis University(伊利诺伊大学厄巴纳-香槟分校; 佐治亚理工学院; 圣路易斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将LLM智能体的上下文策划建模为带删失反馈的容量约束assortment问题,证明最优文件规模上界等结论,发现追加指令或不如选最优子集,还验证无关规则会降低语言模型合规性。

AI 中文摘要

在每次会话开始时,大语言模型(LLM)智能体会加载一个固定的上下文文件,例如 $\texttt{ this http URL }$。该文件中的每个已加载 token 在会话后续的每一轮都会被再次计费,且随着文件规模增大,这些文件会降低性能。然而在实际应用中,人类或自动策划者通常会通过追加内容来扩展这些文件。我们将上下文策划问题建模为一个容量约束的 assortment 问题:指令在有限注意力容量下消耗 token;添加一条指令绝不会提升其他指令的合规性,而保留的指令会产生每会话一次的设置成本。我们证明了最优文件规模的上界,该上界与可用候选指令的数量无关,且追加每条具有正独立价值的指令,其净价值可能比选择最优子集差得多。当 token 价格被低估时,token 预算也会限制损失。随后我们研究了可从过往会话中学习到什么,以及这些信息如何指导添加或移除指令的决策。反馈本质上是删失的:已加载指令的益处和危害是可观测的,而缺失的指令仅在其缺失造成危害时才会产生反馈。在该场景下,我们证明删除智能体忽略的指令可能会不可避免地移除有用的指令。我们确定了在添加一条指令前应收集多少证据,此外,当人类审核者每周期只能检查有限数量的编辑时,我们对遗憾进行了界定。实证实验进一步表明,从真实上下文文件中提取的无关规则会降低语言模型的合规性。

英文摘要

At the start of every session, LLM agents load a fixed context file, such as $\texttt{AGENTS.md}$. Each loaded token in the file is charged again in every later round of the session, and these files can degrade performance as they grow in size. However, in practice, human or automated curators usually grow these files by appending. We formulate context curation as a capacitated assortment problem. Instructions consume tokens under a finite attention capacity; adding an instruction never raises the compliance of the others, while retained instructions incur a per-session setup cost. We prove an upper bound on the optimal file size, regardless of the number of available candidate instructions, and that appending every instruction with positive standalone value can be arbitrarily worse in net value than selecting an optimal subset. A token budget also limits the loss when the token price is underestimated. We then examine what can be learned from past sessions and how this information can guide decisions to add or remove instructions. Feedback is inherently censored: the benefits and harms of loaded instructions are observable, whereas missing instructions generate feedback only when their absence causes harm. In this setting, we show that deleting instructions ignored by agents can inevitably remove helpful ones. We characterize how much evidence should be collected before adding an instruction. Besides, we bound regret when human reviewers can inspect only a limited number of edits per period. Empirical experiments further show that irrelevant rules drawn from real context files reduce language-model compliance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑