发表机构
Università di Bologna; Sant’Anna School of Advanced Studies(博洛尼亚大学; 圣安娜高等研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CodeGraph,一种利用代码专用大语言模型和Wikidata锚定构建的源代码开放分类知识图谱,包含约1.58亿节点和10亿条边,覆盖14种编程语言,并引入质量保证协议验证标注精度。
AI 中文摘要
公共软件仓库(如GitHub和Software Heritage Archive)存储了数十亿个文件,然而从中提取隐含的工程知识——即它们实现的算法、遵循的范式、实例化的模式以及服务的应用领域——仍然具有挑战性,因为现有工具局限于语法和词元级分析。我们提出了一种使用代码专用大型语言模型构建源代码开放分类语义标注的流水线。提取的实体通过三阶段链接过程锚定到Wikidata:确定性SPARQL阶段处理无歧义实体,深度研究智能体解决剩余的长尾问题,层级汇总阶段导入每个已解析Wikidata标识符的父类闭包。生成的标注被具体化为源代码特定的开放分类知识图谱。我们进一步引入了一个校准的质量保证协议,通过结合小型人工金标准集和LLM-as-a-judge过滤器来量化标注精度。我们将流水线应用于Stack-Edu语料库的1.67亿个文件,创建了首个已知的大规模源代码开放分类知识图谱。我们的图谱名为CodeGraph,包含约1.58亿个节点,其中包括约1.45亿个文件、约63,000个提取的概念实体(如算法、范式、设计模式和应用程序领域)以及约19,800个锚定的Wikidata实体。此外,CodeGraph拥有约10亿条类型化边,将文件连接到各自的概念,将这些概念链接到其锚定的Wikidata标识符,并将它们关联到父类别,覆盖14种编程语言。
英文摘要
Public software repositories, like GitHub and Software Heritage Archive, store billions of files, yet extracting their implicit engineering knowledge ---i.e., the algorithms they implement, the paradigms they follow, the patterns they instantiate, and the application domains they serve--- remains challenging, as current tools are constrained to syntactic and token-level analysis. We present a pipeline for building an open-taxonomy semantic annotation of source code using a code-specialised Large Language Model. The extracted entities are grounded in Wikidata through a three-stage linking procedure: a deterministic SPARQL stage handles unambiguous entities, a Deep Research Agent resolves the residual long tail, and a hierarchy-rollup stage imports the parent-of closure of each resolved Wikidata identifier. The resulting annotations are materialised as a source-code-specific open-taxonomy knowledge graph. We further introduce a calibrated quality-assurance protocol that quantifies annotation precision by combining a small human gold set with an LLM-as-a-judge filter. We applied our pipeline to the 167 million files of the Stack-Edu corpus, creating the first known large-scale open-taxonomy knowledge graph for source code. Our graph, named CodeGraph, contains approximately 158 million nodes, which include around 145 million files, about 63,000 extracted concept entities (such as algorithms, paradigms, design patterns, and application domains), and roughly 19,800 grounded Wikidata entities. Furthermore, CodeGraph features approximately 1 billion typed edges that connect files to their respective concepts, link these concepts to their grounded Wikidata identifiers, and relate them to their parent categories, covering 14 programming languages.
CommentsAccepted at CIKM 2026