知识图谱内容嵌入中的高效稠密向量检索
Efficient Dense Vector Search within Knowledge Graph Content Embeddings
- Institute of Visual Computing, Graz University of Technology(格拉茨工业大学视觉计算研究所)
- Department of Information Engineering, University of Padua(帕多瓦大学信息工程系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出QUIVER扩展QLever以原生支持RDF知识图谱中的稠密向量检索,通过张量函数注册、词汇表时解析和虚拟SERVICE索引三项优化,在BSBM和DBpedia基准上实现高达355倍加速,并支持秒级跨模态向量连接。
AI中文摘要:
知识图谱是当今知识基础设施的核心组成部分,支持推理并将知识系统锚定到可验证的事实上。RDF存储和SPARQL引擎实现了这一功能,能够对结构化知识执行一系列检索和推理任务。将它们与语言模型(LMs)结合,将检索增强生成(RAG)扩展到神经符号推理,其中结构化查询门控或重新排序生成输出。这一推理路线要求SPARQL评估原生支持稠密嵌入上的张量操作,使得多模态查询和基于学习的相似性排序能够与图结构约束一起表达。这种方法只有在引擎能够高效执行稠密向量检索时才可行。我们提出了QLever-统一索引向量嵌入检索(QUIVER),这是对QLever的扩展,为RDF知识图谱内的稠密向量检索添加了原生支持。它实现了三项优化:引擎级别的张量函数注册、词汇表时解析JSON编码向量,以及一个在查询内部暴露向量索引的虚拟SERVICE。我们提出了两个新基准:一个带有文本嵌入的柏林SPARQL基准(BSBM)扩展,以及一个带有图像嵌入的DBpedia扩展。与基线相比,仅词汇表时解析在BSBM上单类型排序的中位加速比最高达41.9倍,在DBpedia上达20倍;添加近似最近邻索引后,在BSBM上加速比达355倍,在DBpedia上达97.8倍。该索引还使DBpedia上的跨模态向量连接在几秒内可行,而所有非索引配置均超时。
英文摘要:
Knowledge graphs are a core component of today's knowledge infrastructure, supporting reasoning and anchoring knowledge systems to verifiable facts. RDF stores and SPARQL engines fulfill this function, enabling a range of retrieval and inference tasks on structured knowledge. Coupling them with Language Models (LMs) extends RAG toward neurosymbolic reasoning, where structured queries gate or re-rank generative outputs. This line of reasoning requires that SPARQL evaluation natively support tensor operations on dense embeddings, enabling multimodal querying and learned similarity-based ranking to be expressed together with graph-structural constraints. This approach is feasible only if the engine can efficiently perform dense vector search. We present QLever-Unified Indexed Vector Embedding Retrieval (QUIVER), an extension to QLever that adds native support for dense vector retrieval within RDF knowledge graphs. It implements three optimizations: engine-level registration of tensor functions, vocabulary-time parsing of JSON-encoded vectors, and a virtual SERVICE that exposes a vector index inside the query. We propose two new benchmarks: an extension of Berlin Sparql Benchmark (BSBM) with text embeddings and an extension of DBpedia with image embeddings. Against the baselines, vocabulary-time parsing alone yields median speedups of up to 41.9x on BSBM and 20x on DBpedia for single-type ranking; adding an approximate nearest-neighbor index yields speedups of 355x on BSBM and 97.8x on DBpedia. The index further makes cross-modal vector joins on DBpedia feasible in seconds, whereas all non-indexed configurations time out.