arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

表格数据中多属性逻辑依赖与函数依赖的可扩展提取与可视化

Scalable extraction and visualization of multi-attribute logical and functional dependencies in tabular data

Chaithra Umesh, Arvind Lomrore, Neethu D, Kristian Seegel-Schultz, Saptarshi Bej, Olaf Wolkenhauer

arXiv 2610.08287首次发表:更新:

AI 中文总结

针对表格数据中多属性逻辑依赖和函数依赖的可扩展提取与可视化问题,提出LDTool和HLDTool框架,通过超图缩减搜索空间,在模拟和真实数据集上实现高效且可解释的依赖发现。

AI 中文摘要

理解表格数据中属性间的结构关系是机器学习和模式识别的基础。虽然函数依赖(FD)发现已被广泛研究,但逻辑依赖(LD)的可扩展发现,尤其是在属性数量和依赖阶数增加时,仍未被充分探索。这些依赖捕获了成对或多属性之间的非确定性、条件特定的关系。此外,现有方法未提供提取多属性LD和FD的统一框架。为解决这些局限,我们提出了LDTool和HLDTool,用于从表格数据中提取和可视化多属性LD和FD。LDTool将依赖发现扩展到成对关系之外,而HLDTool通过超图引导的搜索空间缩减实现可扩展提取。在三个模拟和十一个真实世界数据集上的实验表明,所提出的框架提取了有意义的LD和FD,同时提高了可扩展性。LDTool在高维特征空间中恢复了与现有FD发现方法相同的FD,且运行时更低,而HLDTool能够在具有数百个特征的数据集中进行依赖发现。该框架提供了依赖结构的可解释可视化,并支持探索性数据分析和合成表格数据定量评估等应用。

英文摘要

Understanding the structural relationships among attributes in tabular data is fundamental to machine learning and pattern recognition. While functional dependency (FD) discovery has been extensively studied, scalable discovery of logical dependencies (LDs), particularly as the number of attributes and dependency order increase, remains underexplored. These dependencies capture non-deterministic, condition-specific relationships among pairwise or multiple attributes. Furthermore, existing approaches do not provide a unified framework for extracting multi-attribute LDs and FDs. To address these limitations, we propose LDTool and HLDTool for extracting and visualizing multi-attribute LDs and FDs from tabular data. LDTool extends dependency discovery beyond pairwise relationships, while HLDTool enables scalable extraction through hypergraph-guided search-space reduction. Experiments on three simulated and eleven real-world datasets demonstrate that the proposed framework extracts meaningful LDs and FDs while improving scalability. LDTool recovers the same FDs as existing FD discovery methods with lower runtime in high-dimensional feature spaces, whereas HLDTool enables dependency discovery in datasets with hundreds of features. The proposed framework provides interpretable visualizations of dependency structures and supports applications in exploratory data analysis and the quantitative evaluation of synthetic tabular data.

Comments31 pages, 4 figures, submitted to Pattern Recognition Journal

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑