arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先剪枝,后决策:基于JEVDB的可扩展语义查询处理

Prune First, Decide Fast: Scalable Semantic Query Processing with JEVDB

Zhengle Wang, Hanxu Yan, Fuheng Zhao, Chunwei Liu

arXiv 2610.02046首次发表:更新:

发表机构

Purdue Data & AI System Lab (PDAIS); University of Utah(普渡数据与人工智能系统实验室; 犹他大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

JEVDB通过类型化决策模型和语义布隆过滤器减少语义查询延迟与成本,在SemBench和Shelob上实现低延迟、低成本和高F1,同时保持答案质量。

AI 中文摘要

语义数据库系统通过基础模型推理扩展SQL以处理非结构化数据,但现有引擎严重依赖自回归大型语言模型进行离散关系决策,导致高延迟和高成本。我们提出JEVDB,一个可扩展的语义数据库系统,采用快速、类型化的决策模型处理语义过滤、连接、分类和排序,同时选择性地将不确定情况升级至生成式大型语言模型。为减少语义连接工作量,JEVDB将基于关系结构的精确Yannakakis风格半连接约简与语义布隆过滤器(SBFs)相结合,后者利用注册的必要条件在潜在语义边之间筛选候选对。我们在SemBench和Shelob(一个源自TPC-DS的语义连接工作负载)上评估JEVDB。在SemBench上,JEVDB-Flash在所有21个评估查询中实现最低延迟,并在19个查询中实现最低成本,同时保持有竞争力的答案质量。在Shelob上,连接规模扩展至540K候选对,JEVDB以95.7%-97.5%的平均F1完成所有查询。SBF筛选在语义评估前移除87.4%的候选对,可复用的条件索引评分进一步将推理模型升级减少55.2%。交互式查询模拟器、源代码和基准可在https://此URL获取。

英文摘要

Semantic database systems extend SQL with foundation-model inference over unstructured data, but current engines rely heavily on autoregressive LLMs for discrete relational decisions, creating high latency and monetary cost. We present JEVDB, a scalable semantic database system that uses fast, typed decision models for semantic filters, joins, classification, and ranking, while selectively escalating uncertain cases to generative LLMs. To reduce semantic-join work, JEVDB combines exact Yannakakis-style semijoin reduction over relational structure with Semantic Bloom Filters (SBFs), which use registered necessary conditions to screen candidates across latent semantic edges. We evaluate JEVDB on SemBench and Shelob, a TPC-DS-derived semantic-join workload. On SemBench, JEVDB-Flash achieves the lowest latency on all 21 evaluated queries and the lowest cost on 19, while maintaining competitive answer quality. On Shelob, where joins scale to 540K candidate pairs, JEVDB completes all queries with 95.7%-97.5% mean F1. SBF screening removes 87.4% of candidate pairs before semantic evaluation, and reusable condition-index scoring further reduces reasoning-model escalations by 55.2%. An interactive query simulator, source code, and benchmarks are available at https://jevdb.org.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑