arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

召回陷阱:固定预算代码上下文下,最大化召回率的检索器配置会降低问题解决效率

The Recall Trap: A Recall-Maximizing Retriever Configuration Reduces Issue Resolution in Fixed-Budget Code Context

Alexander Adkins, Teimuraz Trapaidze

arXiv 2608.14838首次发表:更新:

AI 中文总结

该研究发现,最大化召回率的检索器配置会降低代码修复的问题解决率,建议固定预算下不要按文件硬去重,应针对任务而非检索指标优化打包策略。

AI 中文摘要

代码助手的检索组件通常针对检索指标进行调优:采用提升召回率@k的配置,并假设下游任务的成功率会随之提升。我们在代码修复任务中开展了一项受控案例研究,这并非新现象,而是已知的相关性-多样性与目标不匹配权衡(Levy等人,2025)的已部署标记、执行分级实例。在SWE-bench Verified上,我们将检索器的命中结果注入固定为12个槽位的上下文包,不使用搜索工具,并在其他完全相同的栈中切换一个标记(单文件去重)。该标记对应更高召回率的配置(启用时提供的包中包含目标文件的比例为0.878,禁用时为0.806),但禁用该标记、以文件广度换取文件内深度,会提升单次解决率:gpt-5.6-sol提升7.6个百分点(从39.2%升至46.8%,n=500,McNemar精确检验p=0.0003),且有可预注册的开放权重复现结果,任何评审者均可重新运行(Qwen3.6-27B,提升3.6个百分点,n=499,p=0.0133);两者均通过按仓库聚类的推理验证。该增益与文件内锚点的数量相关,随机块对照实验否定了argmax选择的人为偏差。我们确定了该效应的适用范围:在词汇BM25检索器上效应反转(降低3.2个百分点,存在显著的跨范式交互),在无限制读取智能体下未检测到(效力不足的零假设),在四种语言(SWE-PolyBench,N=617)中效应为正但不显著(提升2.6个百分点,p=0.056),这是已确定的边界而非已确认的扩展。操作层面,在严格的固定预算下:不要按文件进行硬去重,应针对任务而非该标记所调优的指标进行A/B打包策略测试。

英文摘要

Retrieval components for code assistants are tuned against retrieval metrics: a configuration that raises recall@k is adopted, and downstream task success is assumed to follow. We report a controlled case study in code repair, not a new phenomenon but a deployed-flag, execution-graded instance of the known relevance-diversity and objective-mismatch tradeoff (Levy et al., 2025). On SWE-bench Verified we inject a retriever's hits as a fixed 12-slot context pack with no search tools and toggle one flag (one-chunk-per-file deduplication) on an otherwise identical stack. The flag is the higher-recall configuration (gold file present in 0.878 of served packs against 0.806 disabled), yet disabling it, trading file breadth for within-file depth, raises the single-shot resolve rate: gpt-5.6-sol +7.6pp (39.2% to 46.8%, n=500, McNemar exact p=0.0003), and a pre-registered open-weights replication any reviewer can re-run (Qwen3.6-27B, +3.6pp, n=499, p=0.0133); both survive repository-clustered inference. The gain tracks within-file anchor dose, and a random-chunk control refutes an argmax-selection artifact. We map where it holds: it reverses on a lexical BM25 retriever (-3.2pp, significant cross-paradigm interaction), is not detected under unrestricted-Read agents (a powered null), and across four languages (SWE-PolyBench, N=617) is positive but not significant (+2.6pp, p=0.056), a mapped boundary rather than a confirmed extension. Operationally, at a tight fixed budget: do not hard-deduplicate by file, and A/B packing policies against the task, not the metric the flag was tuned to.

Comments24 pages, 2 figures. Reproducibility artifact: Zenodo DOI 10.5281/zenodo.21879550

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑