arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于部分依赖图的高效上下文检索用于仓库级代码生成

Effective and Efficient Context Retrieval via Partial Dependency Graph for Repository-Level Code Generation

Zhongxin Liu, Zhonghao Jiang, Zhifan Ye, Haoye Wang, Jiakun Liu, Xiaoxue Ren

arXiv 2608.01927首次发表:更新:

发表机构

Zhejiang University; Hangzhou City University; Harbin Institute of Technology(浙江大学; 杭州城市大学; 哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对仓库级代码生成中RAG方法的不足,提出基于部分依赖图的DyRetriever,构建DyCoder并在CoderEval、DevEval上取得显著性能提升且效率更高。

AI 中文摘要

基于大语言模型(LLM)的仓库级代码生成旨在利用软件仓库中的可用上下文生成代码,要求LLM对复杂的代码依赖关系进行推理。由于上下文窗口有限且对仓库特定的理解不足,LLM通常依赖检索增强生成(RAG)来纳入相关代码。早期的RAG方法主要采用基于相似度的检索,往往无法检索到目标函数所依赖的代码片段。近期的研究引入了基于图的检索来建模此类依赖关系,但通常依赖人工设计的规则和静态全局图,导致灵活性有限且构建与维护成本较高。相比之下,人类开发者会通过隐式构建部分依赖图并沿其迭代检查来收集有用的上下文。受此行为启发,我们提出了DyRetriever,一种通过部分依赖图实现的高效上下文检索方法。DyRetriever首先使用LLM选择一组入口点函数,然后沿代码依赖图执行多跳推理。在多跳推理过程中,它利用LLM的语义理解来验证某个函数是否有助于生成目标函数,消除了人工设计的规则,并实现了跨场景的灵活性。DyRetriever不会静态构建全局依赖图,而是按需构建部分图,使用后丢弃,从而降低了构建和维护成本。我们将DyRetriever与基于相似度的代码检索器集成,构建了DyCoder,并在CoderEval和DevEval上对其进行评估。实验结果表明,与现有的基于RAG的方法相比,DyCoder在CoderEval和DevEval上分别实现了25.63%和59.73%的相对Pass@1提升,同时比基于静态依赖图构建的基准方法快7.4倍。

英文摘要

LLM-based repository-level code generation aims to generate code using the context available in a software repository, requiring LLMs to reason over complex code dependencies. Due to limited context windows and insufficient repository-specific understanding, LLMs typically rely on retrieval-augmented generation (RAG) to incorporate relevant code. Early RAG approaches primarily employ similarity-based retrieval, which often fails to retrieve code snippets that the target function depends on. Recent work introduces graph-based retrieval to model such dependencies, but typically relies on manually designed rules and static global graphs, leading to limited flexibility and high construction and maintenance costs. In contrast, human developers collect helpful context by implicitly constructing a partial dependency graph and iteratively inspecting along it. Inspired by this behavior, we propose DyRetriever, an efficient context retrieval method via partial dependency graphs. DyRetriever uses an LLM to first select a set of entry-point functions and then perform multi-hop reasoning along the code dependency graph. During multi-hop reasoning, it uses the LLM's semantic understanding to validate whether a function can help generate the target function, eliminating manually designed rules and enabling flexibility across scenarios. Instead of statically constructing a global dependency graph, DyRetriever builds a partial graph on demand and discards it after use, reducing construction and maintenance costs. We integrate DyRetriever with a similarity-based code retriever to build DyCoder and evaluate it on CoderEval and DevEval. Experimental results show that DyCoder achieves relative Pass@1 improvements of 25.63% and 59.73% on CoderEval and DevEval, respectively, compared with existing RAG-based methods, while being 7.4x faster than baselines based on static dependency graph construction.

CommentsAccepted by ASE 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑