arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.00891cs.IRcs.AI

通过主题增强嵌入提升检索增强生成:一种整合传统NLP技术的混合方法

Enhancing Retrieval-Augmented Generation with Topic-Enriched Embeddings: A Hybrid Approach Integrating Traditional NLP Techniques

  • National University of Tierra del Fuego(泰德菲国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Rodrigo Kataishi

更新

AI总结:

本文提出一种整合传统NLP技术的混合方法,通过主题增强嵌入提升检索增强生成系统的检索精度和语义聚类效果。

AI中文摘要:

检索增强生成(RAG)系统依赖于准确的文档检索来为大型语言模型(LLMs)提供外部知识,但检索质量在主题重叠和主题变化高的语料库中往往会下降。本文提出了一种主题增强嵌入,该嵌入结合了基于术语的信号和主题结构与上下文句子嵌入。该方法结合TF-IDF与主题建模和降维,使用潜在语义分析(LSA)和潜在狄利克雷分配(LDA)来编码潜在的主题组织,并将这些表示与紧凑的上下文编码器(all-MiniLM)融合。通过同时捕捉词级和主题级语义,主题增强嵌入提高了语义聚类,增加了检索精度,并减少了与纯上下文基线相比的计算负担。在法律文本语料库上的实验显示,在聚类一致性和检索指标上均取得一致的提升,表明主题增强嵌入可以作为更可靠知识密集型RAG管道的实用组件。

英文摘要:

Retrieval-augmented generation (RAG) systems rely on accurate document retrieval to ground large language models (LLMs) in external knowledge, yet retrieval quality often degrades in corpora where topics overlap and thematic variation is high. This work proposes topic-enriched embeddings that integrate term-based signals and topic structure with contextual sentence embeddings. The approach combines TF-IDF with topic modeling and dimensionality reduction, using Latent Semantic Analysis (LSA) and Latent Dirichlet Allocation (LDA) to encode latent topical organization, and fuses these representations with a compact contextual encoder (all-MiniLM). By jointly capturing term-level and topic-level semantics, topic-enriched embeddings improve semantic clustering, increase retrieval precision, and reduce computational burden relative to purely contextual baselines. Experiments on a legal-text corpus show consistent gains in clustering coherence and retrieval metrics, suggesting that topic-enriched embeddings can serve as a practical component for more reliable knowledge-intensive RAG pipelines.

↑