arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

搜索引擎的连接索引:针对Lucene段的可剪枝并行半连接

Join Indices for Search Engines: a Prunable Parallel Semijoin over Lucene Segments

Mikhail Khludnev

arXiv 2608.01173首次发表:更新:

AI 中文总结

该研究将关系型系统的Valduriez连接索引技术应用于Lucene段存储,提出可剪枝并行半连接算法,原型在Solr上实现,大幅降低了多产品SKU连接的查询延迟,且优势随负载提升。

AI 中文摘要

连接在搜索引擎中属于二等公民:Lucene现有的查询时连接实现要么在性能上受限,要么在能力上受限,迫使人们在仅作用于单个索引的快速连接与跨独立管理索引的慢速连接之间做出选择。我们将关系型系统中的Valduriez连接索引技术应用于Lucene基于刷写的(LSM风格)段存储:对于每对父段和子段,我们物化一个仅追加的、序对序的连接索引列J[c]=p,避免对外部可变长度键进行任何查询时转换。在此结构之上,我们构建了一种半连接算法,该算法针对每个父段并行计算,阶段间无全局屏障;它在三个层级进行剪枝(段级,其中第一层可通过逐段执行免费获得;先验最小/最大值;以及文档级两阶段确认,采用延迟累积的半读并集),使其可与任意引擎查询组合,而非在后续会被兄弟过滤器丢弃的匹配上浪费计算。作为Apache Solr查询解析器实现的原型,在1000万SKU上连接100万产品的基准测试中,与Solr内置的查询时连接相比,平均查询延迟降低了5.4倍(359.8毫秒 vs. 1934.6毫秒),且该优势随负载单调扩大,在8并发时达到8.3倍:在4个vCPU上,基准峰值为1.18查询/秒后吞吐量下降,而连接索引仍在提升,达到基准最佳值的8.04-6.8倍。

英文摘要

Joins are second-class citizens in search engines: existing query-time join implementations in Lucene are limited either in performance or in capability, forcing a choice between fast joins scoped to a single index and slower joins that span independently managed indices. We carry Valduriez's join-index technique from relational systems to Lucene's flush-based (LSM-style) segment storage: for every pair of a parent and a child segment we materialize an append-only, ordinal-to-ordinal join-index column J[c]=p, avoiding any query-time translation of external variable-length keys. On top of this structure we build a semijoin algorithm that is computed per parent segment, in parallel, without a global barrier between stages; it prunes at three levels (segment-level, the first of which comes free from per-segment execution; a-priori min/max; and document-level two-phase confirmation with a lazily accumulated half-read union) so that it composes with arbitrary engine queries instead of wasting computation on matches that a sibling filter would later discard. A prototype implemented as an Apache Solr query parser, benchmarked on 1M products joined against 10M skus, cuts average query latency 5.4 times (359.8,ms vs. 1934.6,ms) relative to Solr's built-in query-time join, and the advantage widens monotonically with load, reaching 8.3 times at a concurrency of eight: on 4 vCPUs the baseline peaks at 1.18 queries/s and then loses throughput, while the join index is still gaining, at 8.04 - 6.8times the baseline's best.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑