发表机构
Novelcore; University of Peloponnese(诺瓦核心; 伯罗奔尼撒大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Quanta是一个开源Python库,通过加权倒数排名融合统一量化向量搜索、BM25全文检索和知识图谱遍历,并以图作为候选扩展器而非评分器,实现高效混合检索。
AI 中文摘要
一个先进的检索增强生成流水线通常由三或四个独立运行的系统组装而成:一个近似最近邻索引、一个全文检索引擎、一个图数据库和一个关系型文档存储。每个系统都有其自身的部署界面、配置模型和故障模式,而将它们绑定在一起的集成逻辑在每个项目中都要重新编写。在这项工作中,我们提出了Quanta,一个开源Python库,它将基于4比特量化嵌入的稠密向量搜索、BM25全文检索和知识图谱遍历统一在单一检索API之后。Quanta做出了两项设计承诺,这使其区别于现有的混合检索栈。首先,信号通过加权倒数排名融合进行组合,而不是将异构分数归一化到共享范围,我们认为后者是不适定的,因为这种归一化依赖于查询。其次,图是候选扩展器而非相关性评分器:遍历扩大了候选池,新加入的文档在标识符允许列表下由稠密索引重新评分,因此结构邻接决定考虑哪些内容,而内容证据决定其排名。
英文摘要
An advanced retrieval-augmented generation pipeline is typically assembled from three or four independently operated systems: an approximate nearest-neighbour index, a full-text search engine, a graph database, and a relational document store. Each contributes its own deployment surface, configuration model, and failure modes, and the integration logic that binds them is written anew in every project. In this work, we present \textsc{Quanta}, an open-source Python library, which unifies dense vector search over 4-bit quantised embeddings, BM25 full-text retrieval, and knowledge-graph traversal behind a single retrieval API. Quanta makes two design commitments, which distinguish it from existing hybrid retrieval stacks. First, signals are combined by \emph{weighted reciprocal rank fusion} rather than by normalising heterogeneous scores onto a shared range, which we argue is ill-posed because such normalisations are query-dependent. Second, the graph is a \emph{candidate expander and not a relevance scorer}: traversal widens the candidate pool, and the newly admitted documents are re-scored by the dense indexes under an identifier allowlist, so structural adjacency determines what is considered while content evidence determines how it ranks.