发表机构
Hangzhou Innovation Institute, Beihang University(北京航空航天大学杭州创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对RAG检索阶段仅捕捉关联关系的问题,通过建模检索过程的因果结构,提出一种无需训练的注意力式重评分规则,在企业知识库和关键词堆砌语料库上显著提升了检索精度。
AI 中文摘要
检索增强生成(RAG)将大语言模型(LLM)的生成过程建立在检索到的文档基础上,但标准的终端检索阶段——密集向量相似度计算,可选搭配重排序——常常返回与查询共享关键词却不包含所需信息的文档,这种失效模式会随着知识库规模的扩大而加剧。我们将其归因于概念上的缺口:相似度仅捕捉关联关系,而真正重要的文档与查询之间存在因果关联。我们基于赖兴巴赫的共同原因原则,用因果图对终端检索阶段进行建模:查询与检索到的文档共享的关键词构成潜在共同原因A,文档的剩余关键词构成将文档与理想输出关联起来的潜在集合B。由于检索到的文档是对撞体(A -> d <- B),检索本身会在查询与B之间打开一条关联路径,这为一种无需训练的注意力式重评分规则提供了依据:查询嵌入与B的加权质心嵌入之间的余弦相似度。与在知识内容内部建模因果关系的因果增强RAG变体不同,我们的图模型对检索过程本身的因果结构进行建模。在包含471个文档的真实企业知识库上,该方法将一条相关指南的排名从第6位提升至前3位;在重现关键词堆砌机制的受控诊断语料库上,它将平均目标排名从2.88提升至1.25,而训练过的交叉编码器重排序器几乎没有帮助(仅提升至2.63)。相反,在三个BEIR基准测试中,该方法的得分低于相似度基线,这明确了其适用边界:该方法适用于专有知识库规模扩大的关键词堆砌机制,可作为神经重排序器的补充;一个语料库级别的校准门以≥95%的可靠性选择正确的机制。一个完全本地的测试床证明了其可部署性。
英文摘要
Retrieval-Augmented Generation (RAG) grounds LLM generation on retrieved documents, but the standard terminal retrieval stage--dense-vector similarity, optionally followed by reranking--often returns documents that share keywords with the query without containing the needed information, a failure mode that grows with the knowledge base. We trace it to a conceptual gap: similarity captures only associational relations, whereas the documents that matter are linked to the query causally. We model the terminal retrieval stage with a causal graph grounded in Reichenbach's common cause principle: the keywords shared by the query and a retrieved document form a latent common cause A, and the document's residual keywords form a latent set B linking the document to the ideal output. Since a retrieved document is a collider (A -> d <- B), retrieval itself opens an associational path between the query and B, which licenses a training-free, attention-style re-scoring rule: the cosine similarity between the query embedding and the weighted centroid embedding of B. Unlike causality-enhanced RAG variants that model causal relations inside the knowledge content, our graph models the causal structure of the retrieval process itself. On a real 471-document enterprise knowledge base, the method promotes a relevant guideline from rank 6 to the top 3; on a controlled diagnostic corpus reproducing the keyword-stuffing regime, it improves the mean target rank from 2.88 to 1.25, while a trained cross-encoder reranker barely helps (2.63). Conversely, on three BEIR benchmarks the score underperforms the similarity baseline, delineating the applicability boundary: the method guards the keyword-stuffing regime of growing proprietary knowledge bases and complements neural rerankers; a corpus-level calibration gate selects the correct regime with >= 95% reliability. A fully local testbed demonstrates deployability.
Comments16 pages, 2 figures, 4 tables, 1 algorithm. Code available at https://github.com/Silk-Road/causal-rag-rerank