无序中的秩序:高效实现顺序依赖发现算法的技术
Order in Desbordante: Techniques for Efficient Implementation of Order Dependency Discovery Algorithms
- Saint-Petersburg University(圣彼得堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究科学密集型数据剖析中顺序依赖发现问题,通过在C++中重新实现FASTOD和ORDER算法并分析瓶颈、提出改进技术,在Desbordante工具中实验,实现性能大幅提升和内存消耗降低。
AI中文摘要:
科学密集型数据剖析专注于发现和验证数据集中的各种模式。本研究关注一种模式——顺序依赖(OD)的发现,即某些列列表按另一列排序,其在数据库查询优化等方面有用。现有方法仅从算法角度处理,未关注实现。本文研究针对不同OD公理的FASTOD和ORDER算法,用C++重新实现以加速并降低内存消耗,分析瓶颈并提出改进技术。在Desbordante工具中实验表明,重新实现版本性能提升达3倍,应用技术后达10倍,内存消耗降低达2.9倍。
英文摘要:
Science-intensive data profiling focuses on discovery and validation of various patterns in datasets. This study considers discovery of one such pattern - order dependency (OD). Simply put, OD states that some list of columns is ordered according to another one. It is of use for database query optimization, data cleaning and deduplication, anomaly detection, and much more. Existing discovery methods have approached this problem solely from the algorithmic standpoint, without focusing on the implementation side. At the same time, this problem is very computationally intensive, and therefore this part should not be ignored, as it brings ODs closer to industrial use. In this paper, we study two algorithms for OD discovery which target different OD axiomatizations - FASTOD and ORDER. We start by reimplementing these algorithms in C++ in order to speed them up and lower their memory consumption. We then analyze their bottlenecks and propose several techniques which improve their performance even further. To perform evaluation, we have implemented these algorithms inside Desbordante - a science-intensive, high-performance, and open-source data profiling tool developed in C++. Experiments have demonstrated a performance improvement of up to 3x obtained by reimplemented versions, and, with the application of our techniques, up to 10x. Memory consumption has been lowered by up to 2.9x.