发表机构
Huawei Technologies(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有RAG方法在企业知识问答中存在的局限,提出整合向量、目录、图与反思的VDGR-RAG智能体GraphRAG系统,经实验其在召回率和准确率上均优于多种RAG基线。
AI 中文摘要
检索增强生成(RAG)对于企业知识问答(QA)至关重要,尤其在电信等拥有复杂产品文档的领域。然而,现有RAG方法大多忽视了对不同检索优势的整体整合,导致领域路由不准确、分层文档结构利用不足,进而限制了对企业知识的推理能力。为解决这些局限,我们提出VDGR-RAG,它将向量检索、目录驱动推理、图遍历与迭代反思整合到统一框架中,用于准确的企业知识问答。具体而言,VDGR-RAG是一个智能体GraphRAG系统,它首先从文档块构建分层异构知识图($\text{H}^2$KG),以保留分层目录结构与语义关系;随后采用一组可自由组合的原子工具来导航$\text{H}^2$KG,这些工具包括:(1)目录增强路由工具,它使用目录(TOC)结构将用户查询路由到合适的特定领域$\text{H}^2$KG;(2)多路径检索工具,它结合向量搜索、基于TOC的智能体搜索与图搜索,实现全面的知识检索;(3)目录回溯工具,用于纠正知识定位偏差;(4)动态反思工具,它迭代规划下一个检索阶段。我们在涵盖四个无线领域(如节能与故障管理)的企业产品文档上开展了广泛实验,实验结果表明,我们的方法在知识检索召回率和问答准确率两方面均显著优于多种RAG基线方法。
英文摘要
Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex product documentation like telecommunications. However, existing RAG approaches largely overlook the holistic integration of diverse retrieval strengths, leading to inaccurate domain routing, poor utilization of hierarchical document structures, and consequently limited reasoning capabilities over enterprise knowledge. To address these limitations, we present VDGR-RAG, which integrates vector retrieval, directory-driven reasoning, graph traversal, and iterative reflection in a unified framework for accurate enterprise knowledge QA. Specifically, VDGR-RAG is an agentic GraphRAG system that first constructs a Hierarchical Heterogeneous Knowledge Graph ($\text{H}^2$KG) from document chunks to preserve both hierarchical directory structures and semantic relationships, and then employs a set of atomic tools for knowledge retrieval that can be freely composed to navigate the $\text{H}^2$KG: (1) a directory-enhanced routing tool that uses table-of-contents (TOC) structures to route user queries to appropriate domain-specific $\text{H}^2$KGs; (2) a multi-route retrieval tool that combines vector search, TOC-based agentic search, and graph search for comprehensive knowledge retrieval; (3) a directory backtracking tool that corrects knowledge localization biases; and (4) a dynamic reflection tool that iteratively plans the next retrieval phase. We conduct extensive experiments on our enterprise product documents across four wireless domains (e.g., energy saving and fault management). Experimental results demonstrate that our method significantly outperforms a variety of RAG baselines in terms of both knowledge retrieval recall and QA accuracy.