发表机构
National University of Singapore (NUS); Peking University(新加坡国立大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究代码智能体仓库上下文检索问题,引入智能体检索基准,涵盖多种检索任务和样本。评估多种检索方法,发现无单一方法占优,存在校准差距,且记录轨迹有缺失。通过试验表明检索得出的初始上下文有优势,神谕金上下文还有提升空间。
AI 中文摘要
现代代码智能体通常通过最终是否生成正确补丁来评估,但补丁生成依赖于早期的上下文获取阶段,即找到任务所需的仓库文件。我们引入了智能体检索基准,这是一个针对此上游检索问题的文件级基准。样本基于真实编码工作流信号构建,并针对冻结的基础提交仓库进行评估,相关性由智能体接下来所需内容定义,而非直接查询文件语义相似性。该基准涵盖四个正向检索任务及一个评估选择性检索的子集。智能体检索基准包含来自25个仓库的427个样本。我们评估了多种检索方法,没有单一的检索家族占主导地位,不同任务的优胜者差异很大。用反事实控制校准的选择性阈值在自然无金标准例上并未提高选择性成功率,存在校准差距。记录的轨迹在部分样本上也会错过每个金文件。一个受控的种子干预试验发现,与随机非金上下文相比,检索得出的初始上下文能以更少的种子后探索产生更高的文件F1,而神谕金上下文仍有很大提升空间。
英文摘要
Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed for the task. We introduce Agent Retrieval Bench, a file-level benchmark for this upstream retrieval problem. Samples are built from real coding-workflow signals and evaluated against frozen base-commit repositories, with relevance defined by what an agent needs next rather than direct query-file semantic similarity. The benchmark covers four positive-retrieval tasks: code2test, comment2context, trace2code, and edit2ripple; a fifth subset evaluates selective retrieval using natural evidence-backed no-gold cases and counterfactual wrong-repository controls. Agent Retrieval Bench contains 427 samples across 25 repositories: 345 positive examples, 50 natural no-gold examples, and 32 counterfactual controls. The corpus includes 308 base-commit snapshots, 392,000 files, and 7.9 million chunks. We evaluate lexical retrieval, RepoMap, open-source embeddings, selective abstention, and logged agent context selection. No single retrieval family dominates: Qwen3-Embedding-4B has the best sample-weighted MRR on positive samples, Qwen3-Embedding-8B the best Recall@20, and RepoMap the best budgeted context yield at 8K tokens, with task-level winners differing substantially. Selective thresholds calibrated with counterfactual controls do not improve selective success on natural no-gold cases, revealing a calibration gap. Logged trajectories also miss every gold file on 27-35 percent of samples. A controlled seed-intervention pilot finds that retrieval-derived initial context yields higher file F1 with less post-seed exploration than random non-gold context, while oracle gold context shows substantial remaining headroom.