发表机构
Sun Yat-sen University; School of Artificial Intelligence, Sun Yat-sen University; Monash University; CSIRO’s Data61; Zhejiang University; Huawei Cloud Computing Technologies Co., Ltd.(中山大学; 中山大学人工智能学院; 莫纳什大学; 联邦科学与工业研究组织数据61研究所; 浙江大学; 华为云计算技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对仓库级代码生成中LLM输入长度受限及关键令牌易出错问题,提出ACToR框架,通过识别关键令牌触发针对性检索并采用位置感知加权,在RepoExec和CoderEval上较SOTA方法分别提升8.4%、15.4%
AI 中文摘要
仓库级代码生成任务要求合成满足任务需求且与目标仓库上下文一致的代码。由于实际仓库常超出大型语言模型(LLM)的输入长度限制,现有方法普遍采用检索增强生成(RAG)提供仓库特定上下文。尽管现有方法改进了仓库上下文检索,但通常将上下文作为任务级支持,未明确识别生成过程中需要细粒度仓库上下文的关键令牌。在LLM的自回归生成过程中,错误常集中在少量决定性位置:一旦这些令牌生成错误,后续代码可能沿错误语义路径发展,最终导致功能失效,我们将这些位置称为“关键令牌”。本文提出ACToR,一种面向仓库级代码生成的自适应关键令牌感知检索框架,ACToR在生成过程中识别关键令牌并按需触发针对性检索,为这些决定性位置提供仓库上下文。此外,我们为密集检索器设计了位置感知加权方法,优先选择对生成更具信息性的上下文。我们在两个代表性仓库级基准RepoExec和CoderEval上评估ACToR,实验结果显示ACToR始终优于现有最优方法,在RepoExec上实现相对提升8.4%,在CoderEval上实现相对提升15.4%。除性能提升外,我们系统量化了关键令牌的影响,揭示其在主要生成失败中的核心作用,强调了针对性检索策略的必要性。我们在该httpsURL提供代码和数据。
英文摘要
The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining consistent with the target repository context. Since real-world repositories often exceed the input length limits of LLMs, existing approaches commonly adopt retrieval-augmented generation (RAG) to provide repository-specific context. Despite improving repository-context retrieval, existing methods typically provide context as task-level support, without explicitly identifying the critical tokens that require fine-grained repository context during generation. During the autoregressive generation process of LLMs, errors often concentrate at a small number of decisive positions: once such tokens are generated incorrectly, subsequent code may follow an incorrect semantic path and eventually lead to functional failure. We refer to these positions as "critical tokens". In this paper, we propose ACToR, an adaptive critical token-aware retrieval framework for repository-level code generation. ACToR identifies critical tokens during generation and triggers targeted retrieval on demand to provide repository context at these decisive positions. In addition, we design a position-aware weighting method for dense retrievers to prioritize context that is more informative for generation. We evaluate ACToR on two representative repository-level benchmarks, RepoExec and CoderEval. Experimental results show that ACToR consistently outperforms state-of-the-art methods, achieving relative improvements of 8.4% on RepoExec and 15.4% on CoderEval. Beyond performance gains, we systematically quantify the impact of critical tokens, revealing their central role in major generation failures and highlighting the necessity of targeted retrieval strategies. We provide the code and data at https://github.com/DeepSoftwareAnalytics/ACToR.
CommentsUnder review