发表机构
CloudWalk, Inc.(云从科技集团股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多智能体和内存增强大语言模型中协调内容占用任务预算的问题,提出RCWT测量协议,通过改变协调内容等控制变量进行实验,发现模型表现受影响,该测试是上下文分配预算测量原语,但非多智能体收益等完整理论。
AI 中文摘要
多智能体和内存增强的大语言模型系统通常将协调内容、共享状态、先前讨论、工具输出、摘要和角色指令放在用于当前任务的相同有限提示中。这产生了一个实际分配问题:在固定上下文预算下组装调用时,花在协调上的每个令牌都无法用于任务指令或证据。我们引入了圆桌上下文窗口测试(RCWT),这是一种用于测量此任务预算位移效应的受控协议。RCWT在控制总预算、位置顺序、任务家族和评分的同时改变协调内容。在主要的上下文相关召回任务中,当窗口大小\(W = 4096\)时,三种商业模型通过适度开销保持接近基线,然后一旦残留参考证据降至几百个令牌就会急剧下降。窗口缩放摘要与特定任务的剩余预算解释一致,而不是固定百分比阈值,但我们将此视为描述性证据而非普遍规律。为了测试当任务证据保持完整时固定预算悬崖是否仍然存在,我们添加了一个完整任务消融:在协调令牌通过扩展总提示长度增加时,完整任务/参考块保持存在。在该设置中,所有测试调用在GPT - 4.1 - mini、Claude Haiku 4.5和Gemini 2.5 Flash上,高达95%的协调率下都能正确返回每个评分字段。这种消融缩小了断言范围:主要的RCWT悬崖最好理解为任务预算位移,而不是证明仅协调量就会在原始开放式任务中导致语义干扰。因此,RCWT是上下文分配预算的测量原语,而不是多智能体收益或会话级协调的完整理论。
英文摘要
Multi-agent and memory-augmented LLM systems often place coordination content, shared state, prior discussion, tool outputs, summaries, and role instructions, inside the same finite prompt used for the current task. This creates a practical allocation problem: every token spent on coordination is unavailable to task instructions or evidence when a call is assembled under a fixed context budget. We introduce the Roundtable Context Window Test (RCWT), a controlled protocol for measuring this task-budget displacement effect. RCWT varies coordination content while controlling total budget, position order, task family, and scoring. In the main context-dependent recall task at $W=4096$, three commercial models remain near baseline through moderate overhead and then degrade sharply once residual reference evidence falls to a few hundred tokens. Window-scaling summaries are consistent with a task-specific residual-budget interpretation rather than a fixed percentage threshold, but we treat this as descriptive evidence rather than a universal law. To test whether the fixed-budget cliff persists when task evidence remains intact, we add an intact-task ablation: the full task/reference block is kept present while coordination tokens increase by expanding total prompt length. In that setting, all tested calls return every scored field correctly across GPT-4.1-mini, Claude Haiku 4.5, and Gemini 2.5 Flash up to a 95\% coordination ratio. This ablation narrows the claim: the main RCWT cliff is best read as task-budget displacement, not as proof that coordination volume alone causes semantic interference in the original open-ended task. RCWT is therefore a measurement primitive for context-allocation budgeting, not a complete theory of multi-agent benefit or session-level coordination.
Comments10 pages, 1 figure