发表机构
Yildiz Technical University; Intellica Business Intelligence(伊兹丁技术大学; 英特利卡商业智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究在SWE-QA基准上对比语义搜索与深度智能体搜索的代码问答效果,发现语义搜索正确率更高、成本更低,深度智能体搜索存在交接环节的新失败问题。
AI 中文摘要
代码智能体的大量精力都耗费在简单定位仓库内的正确代码上。当前实践主要有两种方法:语义搜索中,智能体从预先构建的仓库向量索引中检索代码块;深度智能体搜索(也被称为子智能体grep搜索)中,规划智能体将探索任务委托给在独立上下文窗口中工作的单独子智能体,且仅返回浓缩结果。第二种设计被视为良好的上下文工程实践,目的是保护主智能体免受上下文污染(又称上下文衰减,即上下文窗口中积累无关材料时发生的准确率损失)。近期的代码智能体(如Claude Code、Codex、Antigravity等)已快速采用该设计,但几乎没有证据表明它能产出更好的答案。我们在仓库级代码问答基准SWE-QA上对比了这两种方法:语义搜索正确回答了65.2%的问题,而深度智能体搜索的正确率为46.2%,且语义搜索给出每个正确答案的成本不到深度智能体搜索的一半。为解释这一差距,我们将每次失败运行编码为失败模式分类体系,该分类体系显示深度智能体搜索并未消除失败,反而引入了一类新失败:其最大的失败份额(41.8%)发生在规划器与其子智能体的交接环节,且这些失败通常是静默的,最终表现为流畅且自信的错误答案。深度智能体搜索解决了一个实际问题,目前是许多代码智能体的首选设计。但我们的结果表明,它提供的保护并非免费,对于可索引仓库的只读问题,检索是更优且成本更低的选择。
英文摘要
Code agents spend much of their effort simply locating the right code inside a repository. Two approaches dominate current practice. In Semantic Search, the agent retrieves code blocks from a vector index built from the repository in advance. In Deep Agentic Search (also known as grep-search by subagent), a planning agent delegates the exploration to a separate subagent that works in an isolated context window and returns only a condensed result. The second design, which is considered good context engineering practice, exists to protect the main agent from context pollution (also known as context rot), the loss of accuracy that occurs as unrelated material accumulates in the context window. Recent code agents (such as Claude Code, Codex, Antigravity, etc) have adopted it quickly, but there is little evidence on whether it produces better answers. We compare the two approaches on SWE-QA, a benchmark for repository-level code question answering. Semantic search answered 65.2% of questions correctly against 46.2% for deep agentic search, and it produced each correct answer at less than half the cost. To explain the gap, we then coded every failed run into a taxonomy of failure modes. The taxonomy shows that deep agentic search did not remove failures but introduced a new class of them: the single largest share of its failures, 41.8%, occurred at the hand-off between the planner and its sub-agent, and these were usually silent, ending in a fluent and confident answer that was wrong. Deep agentic search addresses a real problem and is now the preferred design in many code agents. However, our results show that the protection it offers may not be free, and that for read-only questions over a repository that can be indexed, retrieval was the stronger and cheaper option.
Comments41 pages, 21 figures, 6 tables. Under review at a journal