arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02321cs.DBcs.AIcs.LGcs.PF

FastGFDs:基于Desbordante的高效图函数依赖验证

FastGFDs: Efficient Validation of Graph Functional Dependencies with Desbordante

发表机构圣彼得堡大学 · Universe Data
查看机构详情
  • Saint-Petersburg University(圣彼得堡大学)
  • Universe Data

机构由 AI 辅助整理,请以论文原文为准。

Anton Chernikov, Yurii Litvinov, Kirill Smirnov, George Chernishev

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出FastGFDs算法,基于Desbordante实现GFD验证,针对消费级PC优化,较并行方案性能提升2.6倍、内存消耗降低5倍,是低端单节点环境GFD验证高效算法的首个公开实现。

中文摘要 AI 辅助

图函数依赖(Graph Functional Dependencies,GFD)是一种新近提出的概念,旨在同时捕获图的拓扑结构及属性间的函数依赖关系。验证给定GFD是否在特定图上成立的过程称为GFD验证,该问题计算开销极大,其中定位合适子图的操作约占总运行时间的99%。该概念的提出者最初提出了一种并行方案(算法),专门针对高性能服务器集群设计。本研究的目标是让更广泛的公众能够使用GFD验证,使其可在消费级个人电脑上运行。我们的初步实验表明,现有算法可能并不适合此类场景。因此,我们提出了FastGFDs,这是一种采用新近开发的图匹配技术的GFD验证算法。与并行方案不同,它是串行算法,在整个图上运行。其创新之处在于采用了核心优先分解(Core-First Decomposition)和紧凑路径索引(Compact Path Index,CPI)。我们将其与朴素串行算法及并行方案进行对比,评估了运行时间和内存消耗。本研究是在低端单节点环境中设计高效GFD验证算法的第一步,我们还提供了针对大型数据图进行GFD验证的开源实现。据我们所知,这是该问题算法的唯一公开可用实现,该实现基于Desbordante开发,Desbordante是一款面向科学密集型任务的开源高性能数据探查工具。最后,我们在一个真实图上的实验表明,与并行方案相比,FastGFDs的性能提升最高可达3倍(平均提升2.6倍),采用新的子图匹配算法还使内存消耗降低了5倍。

英文摘要

Graph functional dependencies (GFD) are a recently-developed concept aimed at capturing both topological structures in graphs and functional dependencies between attributes. The process of verifying whether a given GFD holds over a particular graph is referred to as GFD validation. In this very computationally expensive problem, locating suitable subgraphs accounts for about 99% of the total run time. The concept's authors originally proposed a parallel scheme (algorithm), targeting specifically clusters of high-performance servers. The goal of this study is to open GFD validation to a broader public by making it possible to run it on a consumer class PC. Our initial experiments demonstrated that the existing algorithm may not be optimal for these purposes. Therefore, we propose FastGFDs - a GFD validation algorithm that employs a recently developed graph matching technique. In contrast to the parallel scheme, it is sequential and operates on the entire graph. Its novelty lies in the use of Core-First Decomposition and the Compact Path Index (CPI). We compare it with the naive sequential algorithm and the parallel scheme, evaluating run times and memory consumption. The current study is the first step towards designing an efficient algorithm for GFD validation in low-end single-node environments. We also provide an open-source implementation of GFD validation over large data graphs. To the best of our knowledge, this is the only publicly available implementation of an algorithm for this problem. It is developed in Desbordante - an open-source high-performance data profiler aimed at science-intensive tasks. Finally, our experiments on a real-life graph demonstrated up to three times performance (2.6x on average) improvement over the parallel scheme. Employing the new subgraph matching algorithm also reduced memory consumption by five times.

补充信息

↑