arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02213cs.DBcs.AIcs.DCcs.LGcs.PF

使用Desbordante快速发现包含依赖

Fast Discovery of Inclusion Dependencies with Desbordante

Alexander Smirnov, Anton Chizhov, Ilya Shchuckin, Nikita Bobrov, George Chernishev

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对包含依赖发现的计算成本问题,在开源数据探查器Desbordante中实现两种算法并优化,使Spider提速5倍、Faida提速8倍,解决了现有研究忽略实现层面的问题。

中文摘要 AI 辅助

包含依赖是表属性间的一种关系,可指示可能的主键-外键引用。自动发现包含依赖是学术界和工业界都关注的问题,该问题的核心关注点是发现过程的效率,因为它是一项计算成本高昂的任务。然而,现有研究仅关注算法方面,而忽略了实现层面,同时工程细节与算法细节对于实现良好性能至少同等重要。本文描述了两种包含依赖发现算法(Spider和Faida)的高效实现技术:第一种是经典算法,其思想是许多其他包含依赖发现算法的基础,我们提出了一种高效的并行化技术,该技术在大幅加速算法的同时降低了内存消耗;第二种是最先进的近似算法,我们通过四种优化技术对其进行改进:数据缓冲、启用SIMD的执行、精心选择哈希表以及并行化。为了对我们的技术进行实验评估,我们在Desbordante(一个用C++编写的开源科学密集型数据探查器)中实现了这些算法。对于Spider,我们评估了几种不同的选项;对于Faida,我们证明了所有优化技术均有效。我们还将我们的实现与基于Java的数据探查器Metanome进行了比较。总体而言,我们报告Spider的运行时间最多可提升5倍,Faida的运行时间最多可提升8倍。

英文摘要

Inclusion dependency is a relation between attributes of tables that indicates possible Primary Key-Foreign Key references. Automatic discovery of inclusion dependencies is a relevant problem for both academic and industrial communities. The core concern for this problem is the efficiency of discovery process, since it is a computationally expensive task. However, existing studies only address the algorithmic side, while leaving out the implementation aspect. At the same time, engineering details are at least as important as the algorithmic ones for achieving good performance. In this paper, we describe techniques for efficient implementation of two algorithms for discovery of inclusion dependencies - Spider and Faida. The first one is a classic algorithm whose ideas lie in the foundation of many other inclusion dependency discovery algorithms. We propose an efficient parallelization technique, which greatly speeds up the algorithm while simultaneously reducing its memory consumption. The second one is the state-of-the-art approximate algorithm, which we approach by applying four types of optimizations: data buffering, SIMD-enabled execution, careful hash-table selection and parallelization. In order to experimentally evaluate our techniques, we have implemented these algorithms in Desbordante - an open-source science-intensive data profiler written in C++. For Spider, we have evaluated several different options, and in case of Faida we have demonstrated that all our optimization techniques yield results. We also compared our implementations with Metanome - a Java-based data profiler. Overall, we report up to 5x improvement in terms of run time reduction for Spider and up to 8x for Faida.

发表机构

  • Saint-Petersburg State University(圣彼得堡国立大学)
  • Universe Data

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑