arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01450cs.IR

面向边缘设备检索增强生成的双曲空间实时混合检索

Real-Time Hybrid Retrieval in Hyperbolic Space for Retrieval-Augmented Generation on Edge Devices

Aradhya Chakrabarti

AI总结:

该研究提出一种基于洛伦兹模型双曲空间的混合文档检索系统,通过HyTE-H变换投影词嵌入,结合BM25与洛伦兹相似度评分,在BEIR基准数据集上取得良好效果,可在边缘设备实现RAG实时检索。

AI中文摘要:

本文提出一种完全在双曲几何的洛伦兹模型内运行的混合文档检索系统,用于检索增强生成(RAG)。与局限于欧氏空间的传统稠密检索器不同,该系统通过学习到的HyTE-H变换将预训练词嵌入投影到双曲空间,其指数体积增长特性适配自然语言的层级结构。文档被分割为重叠块,通过洛伦兹嵌入建立索引,并采用两阶段流水线检索:首先应用BM25词汇评分,再使用洛伦兹内积相似度对候选进行重排序。可调参数α将BM25评分与双曲相似度评分融合。该系统在BEIR基准套件的五个数据集(SciFact、NFCorpus、ArguAna、SciDocs、FiQA)上进行评估,仅使用词嵌入、未使用微调神经编码器或交叉注意力重排序器时,分别取得0.654、0.304、0.342、0.150、0.217的NDCG@10分数。该系统支持用户提供文档的实时索引,以及对数万份中等规模文档的资源高效查询,因此双曲检索可在边缘设备上以交互延迟运行。

英文摘要:

This paper presents a hybrid document retrieval system designed for retrieval-augmented generation (RAG) that operates entirely within the Lorentz model of hyperbolic geometry. Unlike conventional dense retrievers confined to Euclidean space, this system projects pretrained word embeddings into hyperbolic space through a learned HyTE-H transformation, whose exponential volume growth suits the hierarchical organization of natural language. Documents are segmented into overlapping chunks, indexed by their Lorentz embeddings, and retrieved through a two-stage pipeline that first applies BM25 lexical scoring, then re-ranks candidates using Lorentzian inner-product similarity. A tunable parameter $α$ blends the BM25 score with the hyperbolic similarity score. The system was evaluated on five datasets from the BEIR benchmark suite, SciFact, NFCorpus, ArguAna, SciDocs, and FiQA, achieving NDCG@10 scores of 0.654, 0.304, 0.342, 0.150, and 0.217 respectively with word embeddings alone, without fine-tuned neural encoders or cross-attention rerankers. The system supports real-time indexing of user-supplied documents and resource-efficient querying over tens of thousands of moderately sized documents, so hyperbolic retrieval can run on edge devices at interactive latencies.

↑