arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FeatLens:面向仓库级代码生成的特征引导动态代码图构建与检索

FeatLens: Feature-Guided Dynamic Code Graph Construction and Retrieval for Repository-Level Code Generation

Xutian Li, Bo Xiong, Yifeng Zhu, Kunze Li, Xianlin Zhao, Runbang Yan, Yanzhen Zou, Lu Zhang, Bing Xie

arXiv 2609.26480首次发表:更新:

发表机构

Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FeatLens通过特征索引和动态种子图构建,结合个性化PageRank推理,实现轻量级仓库级代码依赖检索,在DevEval和EvoCodeBench上取得最优DR@15并大幅降低图规模与令牌开销。

AI 中文摘要

近期的代码生成研究已从孤立的函数补全转向现有代码库中的仓库级生成。为了正确实现目标函数,大语言模型(LLM)必须识别可复用的仓库依赖,例如现有函数、API和跨文件定义。现有的检索方法通过代码相似性搜索、持久化全仓库图或LLM驱动的图探索来提供此类上下文,但往往产生较高的图构建、推理和令牌成本。特征导向方法提供了软件功能性的自然视角,然而它们主要支持需求分解、规划或特征编辑,而非代码依赖检索。本文提出FeatLens,一种面向仓库级代码生成的特征引导动态代码图构建与检索方法。FeatLens构建特征索引,将自然语言特征描述链接到函数级代码实体。给定生成任务时,它从特征索引动态构建特定于任务的种子图,并应用基于个性化PageRank的语义-结构图推理来选择紧凑的推理图。该设计以确定性和轻量级的依赖检索取代了持久化全仓库图的维护和LLM探索。在DevEval和EvoCodeBench上的实验表明,FeatLens在稀疏、密集和图基基线中取得了最佳DR@15(分别为0.501和0.460)。在DevEval生成任务上,它获得了最高的DIR@1,使用DeepSeek-V3.2达到52.91%,使用GPT-5-mini达到53.58%,同时保持有竞争力的Pass@1并生成更短的代码。与最强的图基基线相比,FeatLens将图节点减少61.0%,边减少86.2%,总令牌开销减少45.9%,且在检索过程中不使用任何LLM令牌。

英文摘要

Recent code generation research has moved from isolated function completion toward repository-level generation in existing codebases. To implement a target function correctly, an LLM must identify reusable repository dependencies such as existing functions, APIs, and cross-file definitions. Existing retrieval methods provide such context through code similarity search, persistent whole-repository graphs, or LLM-driven graph exploration, but often incur high graph construction, reasoning, and token costs. Feature-oriented methods offer a natural view of software functionality, yet they mainly support requirement decomposition, planning, or feature editing rather than code dependency retrieval. This paper presents \textbf{FeatLens}, a feature-guided dynamic code graph construction and retrieval approach for repository-level code generation. FeatLens builds a feature index that links natural-language feature descriptions to function-level code entities. Given a generation task, it dynamically constructs a task-specific seed graph from the feature index and applies semantic-structural graph reasoning with personalized PageRank to select a compact reasoning graph. This design replaces persistent whole-repository graph maintenance and LLM exploration with deterministic and lightweight dependency retrieval. Experiments on DevEval and EvoCodeBench show that FeatLens achieves the best DR@15 among sparse, dense, and graph-based baselines (0.501 and 0.460). On DevEval generation, it obtains the highest DIR@1, reaching 52.91\% with DeepSeek-V3.2 and 53.58\% with GPT-5-mini, while maintaining competitive Pass@1 and producing shorter code. Compared with the strongest graph-based baseline, FeatLens reduces graph nodes by 61.0\%, edges by 86.2\%, and total token overhead by 45.9\%, with no LLM tokens used during retrieval.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑