发表机构
University of Maryland, Baltimore County (UMBC)(马里兰大学巴尔的摩县分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SelfGraphRAG框架通过从知识图谱结构生成问答对训练查询条件图检索器,解决新建图谱缺乏标注数据的问题,在多跳问答等任务中提升了检索与推理性能。
AI 中文摘要
检索增强生成(RAG)无需重新训练即可通过融入外部知识提升大语言模型性能,但现有方法常未充分利用知识图谱中编码的关系结构。基于图的RAG可捕捉实体关系,不过监督式图检索通常需要带标注的问答数据,而新建的知识图谱可能无法获取此类数据。我们提出SelfGraphRAG框架,该框架直接从知识图谱结构生成问答对,并用其训练查询条件图检索器。生成的问题捕捉多跳路径与局部邻域,无需人工标注即可提供关系监督。在多跳问答与分类基准上的实验显示,SelfGraphRAG相较于基于嵌入的基线方法,提升了检索精度与下游推理性能。这些结果表明,当标注数据不可用时,知识图谱结构可为训练图检索器提供有用的监督。
英文摘要
Retrieval-augmented generation (RAG) improves large language models by incorporating external knowledge without retraining, but existing methods often underuse the relational structure encoded in knowledge graphs. Graph-based RAG can capture entity relationships, yet supervised graph retrieval typically requires labeled question-answer data that may not be available for newly constructed graphs. We address this limitation with SelfGraphRAG, a framework that generates question-answer pairs directly from knowledge graph structure and uses them to train a query-conditioned graph retriever. The generated questions capture multi-hop paths and local neighborhoods, providing relational supervision without manual annotation. Experiments on multi-hop question answering and classification benchmarks show that SelfGraphRAG improves retrieval precision and downstream reasoning performance over embedding-based baselines. These results suggest that knowledge graph structure can provide useful supervision for training graph retrievers when labeled data are unavailable.
Comments16 pages, 2 figures. Based on M.S. thesis work. Thesis available at https://www.proquest.com/docview/3350071346. Under review. Code available upon request from manas@umbc.edu