arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CodeHID:学习用于生成式代码检索的可寻址分层代码索引

CodeHID: Learning an Addressable Hierarchical Code Index for Generative Code Retrieval

Zhen Li, Yuhong Chen, Wenhao Xu, Xiaodong Li, Hui Li

arXiv 2608.24089首次发表:更新:

AI 中文总结

CodeHID是一种生成式代码检索框架,通过伪邻居引导的文档ID学习与双阶段文档ID生成引导构建分层索引,在CoSQA和ProCQA基准上大幅优于现有检索方法

AI 中文摘要

代码检索模型主要依赖将代码片段视为独立候选的扁平匹配范式,使其区分相似代码候选的能力较弱。生成式检索通过在代码语料库上构建可学习索引,引导检索器更好地理解代码候选的语义组织与寻址方式,为该问题提供了解决方案。然而,在代码检索任务中直接应用生成式检索,可能会在其前缀不对应有意义代码语义区域的标识符空间上操作。本文提出CodeHID,一种将代码检索任务从扁平候选匹配重新表述为粗到细语义地址生成的生成式代码检索框架。CodeHID依赖两个核心组件:其一,伪邻居引导的文档ID学习,通过应用多级残差量化和k近邻伪监督构建全局静态分层索引,确保语义相关代码片段共享前缀,同时保留目标级可分性;其二,双阶段文档ID生成引导,通过结合训练侧的排名增强(使用难负样本和排名蒸馏)与推理侧的候选约束及前缀感知解码,可靠地导航该固定索引。在CoSQA和ProCQA基准上的大量实验表明,CodeHID在多数情况下大幅优于现有稀疏检索、预训练代码模型、密集代码检索和生成式检索基线,在排名第一的检索指标上实现了特别显著的改进。

英文摘要

Code retrieval models have predominantly relied on a flat matching paradigm that treats code snippets as independent candidates, making them less capable of distinguishing similar code candidates. Generative retrieval offers a solution by constructing a learnable index over the code corpus, guiding the retriever to better understand how code candidates are semantically organized and addressed. However, naively applying generative retrieval in the code retrieval task may result in operating over an identifier space whose prefixes do not correspond to meaningful code-semantic regions. In this paper, we propose CodeHID, a generative code retrieval framework that reformulates the code retrieval task from flat candidate matching to coarse-to-fine semantic address generation. CodeHID relies on two core components. First, Pseudo-Neighbor Guided DocID Learning constructs a globally static hierarchical index by applying multi-level residual quantization and $k$-nearest-neighbor pseudo-supervision, ensuring that semantically related code snippets share prefixes while preserving target-level separability. Second, Dual-Phase DocID Generation Guidance reliably navigates this fixed index by combining training-side ranking enhancements, using hard negatives and rank distillation, with inference-side candidate constraints and prefix-aware decoding. Extensive experiments on CoSQA and ProCQA benchmarks demonstrate that CodeHID outperforms existing sparse retrieval, pre-trained code models, dense code retrieval, and generative retrieval baselines by a large margin in most cases, achieving particularly strong improvements in rank-one retrieval metrics.

Comments10 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑