发表机构
Integrallis Software(Integrallis软件公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究编码智能体编辑代码时实际所需上下文,通过固定定位、改变代码表示并评分,发现代码本身自然语言摘要作用小,周围上下文无关紧要,压缩上下文效果好,还发现API推理的噪声底限,且发布相关工具。
AI 中文摘要
现代编码智能体可在其上下文窗口中容纳整个存储库。其大部分读取是浪费的,有趣的问题不是智能体能使用多少上下文,而是它实际需要什么。我们在智能体必须编辑代码时研究此问题。将查找工作位置与对其采取行动分开,通过预言机固定定位,仅改变代码表示方式,并根据SWE-bench Verified上的实际问题解决情况对上下文进行评分。答案非常少。信号存在于正在编辑的代码本身:其自然语言摘要几乎无法回答源能回答的行为问题(在保留存储库中,独立评判下为4/45对27/45),差距在于表示方式而非摘要器——前沿模型的摘要得分与3B模型一样差。周围上下文也几乎无关紧要:在Verified中的每个多文件实例中,在任何数据之前冻结的协议下,将文件其余部分呈现为UML骨架和签名解决的问题并不比直接删除其余部分更多(N = 70,精确McNemar检验p = 0.75)。这是我们预先提出的假设,但未成立。同时,压缩上下文以三分之一的令牌匹配整个文件:解决一个问题需要19K上下文令牌,而非94K。该工具还得出一个该领域应留意的发现:温度为0的API推理在字节相同的运行之间会使约9%的每个实例结果翻转。这是此基准测试(包括我们的)中报告的每个小效应下的噪声底限。我们发布该工具——经过黄金验证的环境、每个实例证明每个参考编辑都可从每个分支的上下文中表达、确定性补丁构建以及我们公布零假设的预先注册假设。
英文摘要
A modern coding agent can hold an entire repository in its context window. Most of its reading is wasted -- and the interesting question is not how much context an agent can use, but what it actually \emph{needs}. We study that question at the moment it matters most: when the agent must \emph{edit} code. Separating \emph{finding} the work site from \emph{acting} on it, we hold localization fixed with an oracle, vary only how the code is represented, and score context against real issue resolution on SWE-bench Verified. The answer is starkly minimal. The signal lives in the code being edited itself: natural-language summaries of it answer almost none of the behavioral questions that the source answers ($4/45$ vs.\ $27/45$, held-out repositories, independent judge), and the gap belongs to the representation, not the summarizer -- a frontier model's summaries score exactly as poorly as a 3B model's. The surrounding context hardly matters either: across every multi-file instance in Verified, under a protocol frozen before any data, rendering a file's remainder as UML skeletons and signatures resolves no more issues than deleting that remainder outright ($N{=}70$, exact McNemar $p{=}0.75$). That was our registered hypothesis, and it failed. Compressed context, meanwhile, matches whole files at a third of the tokens: a resolved issue costs $19$K context tokens, not $94$K. The instrument also yielded a finding the field should keep: temperature-0 API inference flips ${\sim}9\%$ of per-instance outcomes between byte-identical runs. That is a noise floor under every small effect reported on this benchmark, including ours. We release the instrument -- gold-validated environments, per-instance proof that every reference edit is expressible from every arm's context, deterministic patch construction, and pre-registered hypotheses whose nulls we publish.