arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ExecRetrieval:衡量代码嵌入检索中的功能正确性差距

ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval

Aaryan Kapoor, Md Abdullah Al Hafiz Khan

arXiv 2609.01865首次发表:更新:

发表机构

Kennesaw State University(肯尼索州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出 ExecRetrieval 基准,含 939 个 Python 任务及配对的规范实现与有缺陷干扰项,评估代码嵌入检索系统,发现其功能正确性差距显著,exec@1 仅 0.331,多数排名第 1 的缺失项为有缺陷变体。

AI 中文摘要

基于嵌入的代码检索是智能体编程和检索增强型代码生成的核心组件,在此场景下,检索正确代码比检索词汇相似代码更为重要。现有代码检索基准未在搜索池中植入每个查询规范实现的受控、经执行验证的单编辑变体,导致在检索场景中,嵌入能否在功能上区分正确代码与近似克隆但不正确代码的问题未得到解答。解决该问题需要一个基准,其搜索池本身包含相关反事实——与每个规范实现近似相同、经执行验证的有缺陷变体,以便直接测试检索器的排名顺序是否具有功能区分能力,而非主题或身份重叠。我们提出 ExecRetrieval,包含 939 个 Python 任务,每个任务配对一个经执行验证的规范实现和最多四个经执行验证的有缺陷干扰项,每个干扰项由一次机械突变生成,仅做一次针对性编辑;我们评估了 23 种密集嵌入配置以及 BM25,采用提供商原生调用方式,并配对使用 McNemar 检验和查询级自举区间。在搜索池包含近似克隆反事实的情况下,顶级托管系统的 exec@10 达到 1.00,但 exec@1 仅为 0.331;在四个领先系统中,排名第 1 的缺失项 91.5%-99.4% 的时间是配对的有缺陷变体,且在领先系统的查询中,规范实现的得分低于其四个配对干扰项中至少一个的情况占 67%-78%。完整数据集、执行预言机、嵌入矩阵、环境快照和配对统计检验已发布在附录 D 中的 URL 处。

英文摘要

Embedding-based code retrieval is a core component of coding agents and retrieval-augmented code generation, where retrieving correct code matters more than retrieving lexically similar code. Existing code-retrieval benchmarks do not plant controlled, execution-verified single-edit variants of each query's canonical implementation in the search pool, leaving the question of whether embeddings can functionally discriminate correct from near-clone-but-incorrect code unanswered in a retrieval setting. Resolving this requires a benchmark whose search pool itself contains the relevant counterfactuals -- execution-verified buggy variants near-identical to each canonical -- so that a retriever's rank ordering can be directly tested for functional discrimination rather than topical or identity overlap. We introduce ExecRetrieval, 939 Python tasks each paired with one execution-verified canonical implementation and up to four execution-verified buggy distractors, each generated by a mechanical mutation making a single targeted edit, and evaluate 23 dense embedding configurations plus BM25 under provider-native invocation with paired McNemar tests and query-level bootstrap intervals. With near-clone counterfactuals in the pool, the top hosted system reaches exec@10 = 1.00 but only exec@1 = 0.331; rank-1 misses are paired buggy variants 91.5-99.4% of the time across the four leading systems, and the canonical scores below at least one of its four paired distractors in 67-78% of queries on the leading systems. The full dataset, execution oracle, embedding matrices, environment snapshot, and pairwise statistical tests are released at the URL in Appendix D.

CommentsAccepted to EMNLP 2026 (Main Conference). Camera-ready version. 17 pages, 6 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑