让你的数据库管理系统(DBMS)学会匹配字符串:用于通配符连接和过滤的快速通用模式匹配
Teach Your DBMS to LIKE Strings: Fast and General Pattern Matching for Wildcard Joins and Filters
浏览论文内容
中文总结 AI 辅助
针对DBMS处理通配符连接与过滤性能差的问题,提出基于Aho-Corasick的连接算法与代码生成式过滤方法,在基准测试中实现显著加速,可用于现代查询引擎的高性能文本分析。
中文摘要 AI 辅助
如今,现代应用不止存储文本,还需从中提取有价值的洞见,通常依赖带LIKE谓词的通配符查询来提取模式。然而,现代数据库管理系统(DBMS)处理这类通配符操作的性能较差,连接操作采用嵌套循环,过滤操作采用开销高昂的解释性执行。为解决连接操作的问题,我们提出一种基于Aho-Corasick算法的新型连接算法,大幅降低时间复杂度;对于通配符过滤,我们利用代码生成基础设施提升性能,为LIKE谓词生成专用代码,消除逐元组解释模式的开销。实验结果显示,新型通配符连接算法的性能显著优于基准系统DuckDB和Umbra,分别实现最高30.6倍和114.75倍的加速;新型通配符过滤方法同样优于两个基准系统,在以过滤为核心的压力基准测试中实现13.3倍的加速。我们认为这两项技术将在现代查询引擎的高性能文本分析中发挥关键作用。
英文摘要
Nowadays, modern applications do more than just store text -- they need to derive meaningful insights from it. To do that, they usually rely on wildcard queries with LIKE predicate to extract patterns. However, modern database management systems (DBMSs) handle these wildcard operations poorly, resorting to nested loops for joins and expensive interpreted evaluation for filters. To address the former, we propose a new join algorithm based on the Aho-Corasick algorithm, which significantly reduces the time complexity. For wildcard filtering, we leverage the code-generation infrastructure to improve performance: we generate specialized code for the LIKE predicate, eliminating the overhead of interpreting the pattern per tuple. Our experimental results show that the new wildcard join algorithm significantly outperforms both baseline DuckDB and Umbra, achieving speedups of up to 30.6x and 114.75x, respectively. The new wildcard filter approach likewise outperforms both baselines, achieving a speedup of 13.3x in a filter-focused stress benchmark. We believe these two techniques will play key roles for high-performance text analytics in modern query engines.