arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于检索增强生成系统的交叉注意力校准去重

Cross-Attention Calibrated Deduplication for Retrieval-Augmented Generation System

Phuong Le Huy, Nam H. Nguyen, Quan V. Dang

arXiv 2607.24332首次发表:更新:

AI 中文总结

研究RAG系统冗余块问题,提出交叉注意力校准去重方法CACD,由交叉编码器比较、新信息分数和多数投票组成,在SQuAD 1.1验证集测试中,能有效去除冗余块且处理速度快。

AI 中文摘要

检索增强生成(RAG)系统中常见的分块策略往往会产生冗余块,这使得向量数据库变大并减慢检索速度。常用的解决方法是余弦相似度阈值处理,但单个向量会丢失区分真正重复块和仅共享相同主题块所需的细粒度、令牌级细节。我们提出了交叉注意力校准去重(CACD),它使用交叉编码器而非单个池化向量,将每个新块与已保存的内存块池进行检查,保留令牌级细节直至最终比较。CACD由三部分组成:交叉编码器比较本身、衡量块中有多少未被已保存候选块解释的新信息分数(NIS)以及对多个候选块而非单个最佳匹配的多数投票。NIS由交叉编码器的注意力熵计算得出。我们在完整的SQuAD 1.1验证集上,针对五种现有过滤方法、九种分块策略和18种配置测试了CACD。实验中,CACD平均去除9.75%的块,处理每个配置平均用时51.0秒,比最强基线NERExact快约27%,比余弦相似度过滤快约7倍。代码可在指定链接获取。

英文摘要

Common chunking strategies in Retrieval-Augmented Generation (RAG) systems often create redundant chunks. These redundant chunks make the vector database bigger and slow down retrieval. A common fix is cosine-similarity thresholding. This method reduces each chunk to a single vector, then compares vectors using a similarity score. But a single vector can lose the fine-grained, token-level detail needed to tell a true duplicate apart from a chunk that just shares the same topic. We propose Cross-Attention Calibrated Deduplication (CACD). CACD checks each new chunk against an in-memory pool of chunks already kept, using a cross-encoder instead of a single pooled vector. This keeps token-level detail all the way to the final comparison. CACD combines three parts: the cross-encoder comparison itself, a New Information Score (NIS) that measures how much of a chunk is not explained by a candidate already kept, and a majority vote across several candidates rather than a single best match. NIS is calculated from the attention entropy of the cross-encoder. We tested CACD against five existing filtering methods, nine chunking strategies, and 18 configurations, all on the full SQuAD 1.1 validation set. In our experiments, CACD removes 9.75% of chunks on average. This drop rate is close to other semantic-level methods, and much higher than exact-match filters, which barely remove anything. In these experiments, CACD also processes each configuration in 51.0 seconds on average, about 27% faster than the strongest baseline, NERExact (69.6s), and about 7x faster than cosine-similarity filtering (356.7s). These results come from a single dataset, so we present them as an early comparison, not a general claim. Code for the baseline evaluation and for CACD is available at https://github.com/lehuyphuong/rag_bench and https://github.com/lehuyphuong/cacd_dedup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑