arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估源代码库的语义和质量感知检索

Evaluating Semantic and Quality-Aware Retrieval for Source Code Repositories

Marek Horváth, Emília Pietriková

arXiv 2607.09161首次发表:更新:

AI 中文总结

研究针对自然语言查询的源代码库检索问题,结合函数级碎片化、文本和代码嵌入等技术及多种检索模式构建原型系统,用C代码语料库评估,发现语义检索总体最强,自动路由表现良好,质量元数据对特定查询有用。

AI 中文摘要

基于关键字的检索在处理自然语言表达的查询、涉及实现意图和代码质量而非精确令牌的源代码库时存在局限性。本研究评估了一个原型检索系统,该系统结合了函数级碎片化、文本和代码嵌入、ChromaDB向量存储、基于大语言模型的质量元数据以及四种检索模式:语义检索、质量过滤检索、混合检索和自动路由检索。具体评估使用了一个教育用C代码语料库。通过15个手动判断的查询,语义检索的nDCG@5为0.820,Success@5为0.800,MRR为0.644。自动路由器为所有15个查询选择了预期模式。在一次小型人工审核中,12个片段中有9个基于大语言模型得出的质量分数与人工评估相差不超过一分。在报告的查询集中,语义检索总体上是最强的模式,而明确的质量元数据对明确以质量为导向的查询最有用。

英文摘要

Keyword-based retrieval is limited for source-code repositories when queries are expressed in natural language or concern implementation intent and code quality rather than exact tokens. This study evaluates a prototype retrieval system that combines function-level fragmentation, text-and-code embeddings, ChromaDB vector storage, LLM-derived quality metadata, and four retrieval modes: semantic, quality-filtered, hybrid, and automatic routing. The concrete evaluation uses an educational C-code corpus. The full corpus contains 563 anonymized programmer identifiers and 8,951 C files; a reproducible 10% indexed sample contains 56 programmer identifiers, 847 files, and 3,839 fragments. Across 15 manually judged queries, semantic retrieval achieved nDCG@5 of 0.820, Success@5 of 0.800, and MRR of 0.644. The automatic router selected the expected mode for all 15 queries. In a small manual audit, LLM-derived quality scores were within one point of the manual assessment for 9 of 12 fragments. Within the reported query set, semantic retrieval was the strongest overall mode, while explicit quality metadata was most useful for explicitly quality-oriented queries.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑