arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SVD-RAG:通过奇异值分解实现高效的树状组织检索增强生成

SVD-RAG: Efficient Tree-Organized Retrieval-Augmented Generation via Singular Value Decomposition

Zhihui Sun

arXiv 2607.10316首次发表:更新:

AI 中文总结

研究提出SVD-RAG,通过对密集句子嵌入矩阵应用奇异值分解进行层次RAG抽取式摘要。该方法具确定性、成本效益高、内容自适应特点,实验表明其检索质量与RAPTOR相近,构建树速度更快且在多主题基准上性能提升显著。

AI 中文摘要

检索增强生成(RAG)系统通过从外部知识库中检索相关文档来增强大型语言模型。Sarthi等人(2024年)的近期工作引入了RAPTOR,它将文档组织成层次树结构以进行高效检索,但在每个内部节点都需要昂贵的基于大语言模型的抽象摘要,这使得大规模部署成本过高。我们提出了SVD-RAG,这是第一种在密集句子嵌入矩阵上应用奇异值分解(SVD)进行层次RAG中的抽取式摘要的方法。与在稀疏TF-IDF矩阵上运行的经典潜在语义分析不同,SVD-RAG利用现代嵌入模型丰富的语义表示,通过主成分中的能量贡献识别最具信息性的句子。我们的方法具有确定性(与基于大语言模型的摘要不同,SVD对相同输入产生相同结果)、成本效益高(树构建除初始嵌入外无需额外的API调用,减少约85%的令牌消耗)以及内容自适应(能量比阈值tau根据内容复杂性自动调整压缩)。在使用相同语料库、聚类和波束搜索的受控头对头比较中,SVD-RAG在构建树的速度快317倍(0.1秒比31.7秒)的情况下,实现了与使用大语言模型摘要的RAPTOR相差1-5%的检索质量(平均倒数排名0.867对0.875,召回率@1 0.483对0.458)。在具有205个块和跨20个主题变体的100个查询的扩展多主题基准上,SVD-RAG在召回率@1上提高了4.2倍,在平均倒数排名上提高了3.1倍。我们提供了详细的成本分析和参数敏感性研究。我们的实现作为开源Python包发布。

英文摘要

Retrieval-Augmented Generation (RAG) systems enhance large language models by retrieving relevant documents from external knowledge bases. Recent work by Sarthi et al. (2024) introduced RAPTOR, which organizes documents into hierarchical tree structures for efficient retrieval, but requires expensive LLM-based abstractive summarization at each internal node -- making large-scale deployment prohibitively costly. We present SVD-RAG, the first method to apply Singular Value Decomposition (SVD) on dense sentence embedding matrices for extractive summarization in hierarchical RAG. Unlike classical LSA which operates on sparse TF-IDF matrices, SVD-RAG exploits the rich semantic representations of modern embedding models, identifying the most informative sentences through their energy contribution in the principal components. Our approach is (1) deterministic -- unlike LLM-based summarization, SVD produces identical results for the same input; (2) cost-efficient -- tree construction requires no additional API calls beyond the initial embedding, reducing token consumption by ~85%; and (3) content-adaptive -- the energy-ratio threshold tau automatically adjusts compression based on content complexity. In a controlled head-to-head comparison using identical corpora, clustering, and beam search, SVD-RAG achieves retrieval quality within 1-5% of RAPTOR with LLM summarization (MRR 0.867 vs. 0.875, Recall@1 0.483 vs. 0.458) while building the tree 317x faster (0.1s vs. 31.7s). On a scaled multi-topic benchmark with 205 chunks and 100 queries across 20 topic variations, SVD-RAG achieves a 4.2x improvement in Recall@1 and 3.1x improvement in MRR over flat embedding retrieval. We provide a detailed cost analysis and parameter sensitivity study. Our implementation is released as an open-source Python package.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑