DIADA:数据湖中的自动数据组合
DIADA: Automatic Data Composition in Data Lakes
查看机构详情
- Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对数据湖中属性组合依赖人工且缺乏通用性的问题,提出DIADA系统,通过多元依赖性标准挖掘有意义关系,提供低噪声属性子集,惠及多样化下游任务。
中文摘要 AI 辅助
数据湖包含大量分散在众多表中的属性,当这些属性组合在一起时,可为数据分析提供增强的资产。然而,决定哪些属性在有意义的关系中应归为一组,仍然是一项手动的、逐任务的工作。仅通过可连接性进行合并无法保证属性相关性,而针对单一目标选择特征则会丢弃对其他任务有用的属性。为解决这一差距,我们引入了数据组合问题:将碎片化、异构的数据湖组织成有意义的关系,且不依赖于任何特定的分析任务,从而使所得到的组织能够作为多样化下游分析的共同基础。我们提出了DIADA,一个组合系统,它采用多元依赖性作为评估关系有意义性的标准,并通过假设属性之间独立、识别违反该假设的属性集合来近似实现这一标准。为此,我们将属性映射到谓词空间,在包含关系下形成格结构,并挖掘那些其组成成分之间表现出依赖性的谓词集合。我们贡献了一个专门的、可扩展的算法来有效探索这一空间,其性能超越了经典的关系挖掘算法,从而发现否则难以识别的依赖性。我们证明,应用单一的数据组合过程有利于多样化的潜在下游任务。这是因为提供了低噪声、统计相关的属性子集,增加了检测到的模式基于真实关系的置信度,从而防止了大规模环境中的常见建模问题。
英文摘要
Data lakes contain a plethora of attributes scattered across many tables that, when combined, provide enhanced assets for data analysis. Nonetheless, deciding which attributes belong together in meaningful relations remains a manual, per-task effort. Merging by joinability alone provides no guarantees regarding attribute relevance, while selecting features against a single target discards attributes useful to other tasks. To address this gap, we introduce the data composition problem: organizing a fragmented, heterogeneous lake into meaningful relations, agnostic of any particular analytical task so that the resulting organization can serve as a common foundation for diverse downstream analyses. We propose DIADA, a composition system that employs multivariate dependence as the criterion for assessing the meaningfulness of a relation and approximates it by hypothesizing independence among attributes and identifying those sets that violate this hypothesis. To do so, we map the attributes to a predicate space, forming a lattice under inclusion and mining those predicate sets that exhibit dependence among their constituents. We contribute a dedicated and scalable algorithm to effectively explore this space, outscaling classical algorithms for mining relationships, thus discovering dependencies that would otherwise be impractical to identify. We demonstrate that applying a single data composition process benefits diverse potential downstream tasks. This is the result of providing a subset of low-noise, statistically relevant attributes that increases the confidence that detected patterns are grounded in real relationships, thus preventing common modeling issues in large-scale environments.