arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22081cs.SEcs.CLcs.IR

W-RAG:面向异构知识库的企业文档生成的源感知检索方法

W-RAG: Source-Aware Retrieval for Enterprise Document Generation from Heterogeneous Knowledge Bases

  • The University of Texas at Dallas(德克萨斯大学达拉斯分校)
  • DIGITAL MANAGEMENT, LLC (DMI)
  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Hridya Dhulipala, Rajesh Ombase, Michael Wang, Tien N. Nguyen

中文总结 AI 辅助

针对异构知识库下企业文档生成中全局排序导致的上下文失衡问题,提出源感知检索框架W-RAG,引入相关数据集并验证其可提升文档覆盖率与生成质量。

中文摘要 AI 辅助

检索增强生成(RAG)使大型语言模型能在生成过程中融入外部知识,提升事实依据性与领域适应性。然而现有RAG流程假设,从多个知识库检索到的证据可通过单一相似度函数进行全局排序,该假设虽适用于开放域检索,但在企业文档生成场景中失效——企业的异构知识库(如政策、法规、技术文档、部门指南)作用各异,需在生成文档中协同呈现,全局排序常导致上下文失衡,被部分来源主导,进而生成不完整的企业文档。为解决此局限,本文提出W-RAG,这是一种源感知检索框架,采用本体引导检索、各知识库内局部排序及源级权重调节证据构成的方法;同时引入了涵盖多种文档类型与行业领域的检索增强企业文档生成新数据集。实验表明,标准RAG流程在该任务上表现不佳,而W-RAG显著提升了文档覆盖率与生成质量。

英文摘要

Retrieval-Augmented Generation (RAG) enables large language models to incorporate external knowledge during generation, improving factual grounding and domain adaptability. However, existing RAG pipelines assume that evidence retrieved from multiple repositories can be ranked globally using a single similarity function. While suitable for open-domain retrieval, this assumption breaks down in enterprise document generation, where heterogeneous knowledge bases (such as policies, regulations, technical documentation, and departmental guidelines) serve distinct roles and must be jointly represented in the generated document. As a result, global ranking often produces unbalanced context dominated by a subset of sources, leading to incomplete enterprise drafts. To address this limitation, we propose W-RAG, a source-aware retrieval framework that performs ontology-guided retrieval, local ranking within each knowledge base, and source-level weighting to regulate evidence composition. We further introduce a new dataset for retrieval-grounded enterprise document generation spanning multiple document types and industry domains. Experiments show that standard RAG pipelines struggle on this task, while W-RAG significantly improves document coverage and generation quality.

补充信息

↑