发表机构
Vodafone Idea(沃达丰创意公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对企业文档搜索难题,提出结合知识图谱扩展、RRF融合及分块 grounded 评估的混合 RAG 系统 DocuSearch,在电信语料库上较基线取得显著性能提升。
AI 中文摘要
从大型企业文档库中获取准确、有依据的答案是一个难题。仅靠稠密向量检索在混合技术术语、厂商特定缩写或需要跨多个非相邻段落推理的查询上表现往往不佳。DocuSearch 正是为解决这一缺口而构建的——这是一个离线多智能体文档智能系统,在电信网络运营的生产环境中开发并评估。DocuSearch 不依赖单一检索信号,而是整合了三个互补的证据来源:使用 BGE-Large 嵌入在 Qdrant 向量存储上的语义搜索、在 SQLite FTS5 索引上的 BM25 全文搜索,以及来自结构化边表的知识图谱邻域扩展。这三个排名列表通过 Reciprocal Rank Fusion( reciprocal rank fusion,互反排名融合)合并,其中向量搜索的信号权重为 0.50,BM25 为 0.35,知识图谱为 0.15,使用平滑常数 60 来稳定分数。随后,交叉编码器对融合后的列表进行重新排序,平衡因子为 0.65 的最大边际相关性(Maximal Marginal Relevance)会修剪结果以兼顾相关性和多样性。DocuSearch 的独特之处在于其分块评估循环,将每个文本块视为一个小型检索问题:大型语言模型(LLM)决定该块是否需要更多上下文、是否完全回答了查询,以及答案是否基于检索到的文本。无依据的答案不会被返回;系统会转而采用多块合并。在电信语料库上,DocuSearch 达到了 Precision@10(前 10 个结果的精确率)为 0.69,Recall@10(前 10 个结果的召回率)为 0.79,有依据率为 89.6%——相较于仅使用稠密 RAG 的基线,分别提升了 15、16 和 18.4 个百分点。
英文摘要
Getting accurate, grounded answers out of large enterprise document repositories is a difficult problem. Dense vector retrieval alone frequently performs poorly on queries that mix technical terminology, vendor-specific acronyms, or require reasoning across several non-adjacent sections. DocuSearch was built to address exactly this gap - an offline, multi-agent document intelligence system developed and evaluated in a production telecom network operations environment. Rather than relying on a single retrieval signal, DocuSearch pulls together three complementary sources of evidence: semantic search over a Qdrant vector store using BGE-Large embeddings, BM25 full text search over an SQLite FTS5 index, and Knowledge Graph neighbour expansion from a structured edge table. These three ranked lists are merged through Reciprocal Rank Fusion with signal weights of 0.50 for vector search, 0.35 for BM25, and 0.15 for the knowledge graph, using a smoothing constant of 60 to stabilize scores. A cross-encoder then reranks the fused list, and Maximal Marginal Relevance with a balance factor of 0.65 prunes results for relevance and diversity. What makes DocuSearch distinctive is a per-chunk evaluation loop treating each chunk as its own mini-retrieval problem: an LLM decides whether the chunk needs more context, whether it fully answers the query, and whether the answer is grounded in retrieved text. Ungrounded answers are not returned; the system falls back to a multi-chunk merge instead. On a telecom corpus, DocuSearch reaches Precision@10 of 0.69, Recall@10 of 0.79, and a grounding rate of 89.6% - gains of 15, 16, and 18.4 percentage points over a dense-only RAG baseline. Index Terms: retrieval-augmented generation, knowledge graph, reciprocal rank fusion, enterprise document search, agentic evaluation, BM25, cross-encoder reranking, on-premise deployment, LangGraph, telecom AI.