智能体不进行分页:面向大语言模型工具响应的首块选择
Agents Don't Paginate: First-Chunk Selection for LLM Tool Responses
浏览论文内容
中文总结 AI 辅助
该研究针对LLM编码智能体的工具响应首块选择问题,发现提升首块Top1准确率不会系统性提升下游准确率,无参数关键词评分器可提升p₁但不影响下游准确率。
中文摘要 AI 辅助
基于大语言模型(LLM)构建的编码智能体,如Claude Code、Cursor、OpenAI Codex、GitHub Copilot和Aider,接收到的工具响应通常会超过智能体每轮的token预算。标准的解决方法是分页,在生成这些响应的所有协议中都可用;但在我们观察到的来自公共模型上下文协议中间件的会话日志语料库中,没有智能体发起对第二块的请求。智能体读取的是第一块,因此我们研究所需的目标项(黄金项)在第一块中排在首位的频率:即Top1准确率p₁。在受控的离线基准测试中,我们将首块选择视为0/1背包问题,在500个SWE-bench Verified任务上比较6种价值函数,然后通过对5种语言模型进行单轮文件定位探测,测试p₁是否重要(共4800次LLM调用;并非端到端解决率测试)。两个预先注册的假设不成立,这是我们的主要发现。核心发现是否定性的:提高p₁并不会系统性提升下游准确率。每个模型的变化幅度均在3个百分点(p.p.)以内,符号不一致,且没有任何模型表现出显著性;智能体可从块中的任何位置恢复黄金项,因此影响其答案的是黄金项是否被包含在首块中,而非黄金项在首块内的排名。第二个发现:向关键词评分器中添加4个文件元数据信号会使p₁降低4.8个百分点(配对显著性检验,p=0.001)。无参数的关键词评分器确实能将p₁从24.2%的基线提升至35.0%(+10.8个百分点,远高于随机水平;p=3.9×10⁻⁸),在无关键词匹配时回退到工具的原生排序,p₁可进一步提升至35.8%。但根据核心发现,这是排名1的增益,而排名1是不会影响智能体答案的部分:下游准确率并未发生变化。
英文摘要
Coding agents built on large language models (LLMs), such as Claude Code, Cursor, OpenAI Codex, GitHub Copilot, and Aider, receive tool responses that routinely exceed the agent's per-turn token budget. The standard remedy, pagination, is available in every protocol that produced these responses; yet across the corpus of session logs from a public Model Context Protocol middleware we observed no agent-initiated requests for a second chunk. The first chunk is what the agent reads, so we ask how often the gold item (the one the agent needs) is placed first in it: the precision-at-1 rate $p_1$. In a controlled offline benchmark we treat first-chunk selection as a 0/1 knapsack and compare six value functions on 500 SWE-bench Verified tasks, then test whether $p_1$ matters with a single-turn file-localisation probe on five language models (4,800 LLM calls; not an end-to-end resolve-rate test). Two pre-registered hypotheses did not hold and are our main findings. The central one is negative: raising $p_1$ does not systematically raise downstream accuracy. Per-model deltas stay under three percentage points (p.p.), are not consistently signed, and no model is significant; the agent recovers the gold from anywhere in the chunk, so what reaches its answer is first-chunk inclusion, not the gold's rank within it. The second: adding four file-metadata signals to a keyword scorer hurts $p_1$ by 4.8 p.p. (paired significance test, $p = 0.001$). A parameter-free keyword scorer does raise $p_1$, from a 24.2% baseline to 35.0% (+10.8 p.p., far beyond chance; $p = 3.9 \times 10^{-8}$), and to 35.8% with a fallback to the tool's native ordering when no keyword matches. But by our central finding this is a rank-1 gain, and rank-1 is the part that does not reach the agent's answer: downstream accuracy does not move.
发表机构
- SnT, University of Luxembourg(卢森堡大学SnT(卢森堡大学安全、可靠性与信任中心))
机构由 AI 辅助整理,请以论文原文为准。