CARET:仓库级代码补全的无训练测试时扩展
CARET: Training-Free Test-Time Scaling for Repository-Level Code Completion
另 2 家 · 查看机构详情
- The University of Arizona(亚利桑那大学)
- California State University, Long Beach(加州州立大学长滩分校)
- East Carolina University(东卡罗来纳大学)
- Nankai University(南开大学)
- Nanyang Technological University, Singapore(新加坡南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
CARET提出无需训练的测试时扩展方法,通过样本一致性路由检索上下文和反向上下文似然选择生成候选,在仓库级代码补全中显著提升精确匹配率。
中文摘要 AI 辅助
检索增强生成(RAG)主导着仓库级代码补全:它检索跨文件上下文(R),然后解码一次贪心补全(G)。现有工作主要聚焦于检索,并止步于此。我们认为这两个阶段可以共同改进,尤其是生成阶段能从测试时扩展中获益。我们提出CARET,一种无需训练的方法。对于R,CARET利用其自身样本间的一致性在检索上下文之间进行路由,当样本分散时,级联到替代上下文。对于G,它在缓存的提示前缀上采样候选,因此长检索上下文被编码一次,而不是每个样本编码一次。然后通过反向上下文似然选择最终补全:正确的补全使得光标后的代码更可能,因此同一模型通过向前阅读来评估其自身候选。在CrossCodeEval、RepoEval-Line和RepoEval-API上,使用六个代码模型(1.1B到7B,四个家族),CARET在全部18个组合中平均比贪心解码提高10.8个精确匹配点,比自一致性@10提高5.3个点。令牌级计算保持在单次生成的1.45倍左右(实测墙钟时间为1.3-2.9倍,随生成器规模增长)。共同改进检索和生成比单独改进检索产生更准确的代码,且预算接近单次通过。
英文摘要
Retrieval-augmented generation (RAG) dominates repository-level code completion: it retrieves cross-file context (R), then decodes one greedy completion (G). Existing work mainly focuses on retrieval and stops there. We argue both stages can be improved together, with generation in particular gaining from test-time scaling. We present CARET, a training-free method. For R, CARET routes among retrieval contexts using the agreement among its own samples, cascading to an alternative context when the samples scatter. For G, it samples candidates over a cached prompt prefix, so the long retrieved context is encoded once rather than once per sample. It then selects the final completion by reverse-context likelihood: a correct completion makes the code after the cursor more probable, so the same model grades its own candidates by reading ahead. Across CrossCodeEval, RepoEval-Line, and RepoEval-API with six code models (1.1B to 7B, four families), CARET improves exact match in all 18 combinations by 10.8 points on average over greedy decoding and 5.3 over self-consistency@10. Token-level compute stays near 1.45 times one generation (measured wall-clock 1.3-2.9 times, growing with generator size). Improving retrieval and generation together yields more accurate code than improving retrieval alone, at a budget that stays close to a single pass.