发表机构
Berliner Hochschule für Technik; Physikalisch-Technische Bundesanstalt; Einstein Center Digital Future; FernUniversität in Hagen; Johannes Kepler University Linz; Hasso Plattner Institute(柏林应用技术大学; 德国联邦物理技术研究院; 爱因斯坦数字未来中心; 哈根开放大学; 约翰内斯·开普勒林茨大学; 哈索·普拉特纳研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对ISO/IEC 25024和ISO/IEC 5259的DQ指标实施问题,将其分类后提出指南,在开源库dqmeasure中实现20个指标,实验验证其可用于自动化DQ监测。
AI 中文摘要
尽管数据质量(DQ)研究已有数十年历史,但文献中以文本形式定义的准确性、完整性等DQ维度,与通常实现与这些维度不一致的低级检查的DQ工具之间仍存在差距。ISO/IEC 25024和ISO/IEC 5259试图通过为每个DQ维度定义DQ指标来弥合这一差距,但这些DQ指标几乎未被使用,因为标准未说明如何实施它们:例如,语法准确性指标会统计语法准确的值,但未说明如何判定一个值在语法上准确,这只是将问题转移到了另一个层面而未解决问题。因此,当前DQ评估无法基于这些标准开展。在本文中,我们使ISO DQ指标可执行:我们将两个标准的所有数据级指标分为三类:(i)无需数据之外输入的可泛化指标;(ii)参数可从干净参考数据中学习或由专家设定的参数化指标;(iii)需要定性判断且无法自动化的不可泛化指标。对于可自动化的指标,即(i)和(ii)类,我们提出了实施指南,解决了标准中未明确的问题。我们在dqmeasure中实现了这些指南,这是一个包含20个指标的开源库,这些指标从参考数据中学习参数,而非依赖手动定义的规则。我们在真实世界和合成数据集上的实验表明,指标得分随注入错误数量的增加而单调下降,随下游ML性能的下降而同步下降,且与行数呈线性缩放,这使得基于标准的自动化DQ监测成为可能。
英文摘要
Despite decades of data quality (DQ) research, a gap remains between DQ dimensions, such as accuracy or completeness, which the literature defines in textual form, and DQ tools, which typically implement low-level checks that are not aligned with these dimensions. ISO/IEC 25024 and ISO/IEC 5259 attempt to bridge this gap by defining DQ metrics for each dimension. However, these DQ metrics are hardly used, because the standards leave open how to implement them: for example, the metric for syntactic accuracy counts syntactically accurate values, but does not state how to decide that a value is syntactically accurate. This simply moves the problem to another level without solving it. As a result, DQ assessment currently cannot build on the standards. In this paper, we make the ISO DQ metrics executable. We classify all data-level metrics of both standards into (i) generalizable metrics that need no input beyond the data, (ii) parameterized metrics whose parameters can be learned from clean reference data or set by an expert, and (iii) non-generalizable metrics that need qualitative judgment and cannot be automated. For the metrics that can be automated, i.e., categories (i) and (ii), we propose implementation guidelines that resolve what the standards leave open. We realize the guidelines in dqmeasure, an open-source library of 20 metrics that learns these parameters from reference data instead of relying on manually defined rules. Our experiments on real-world and synthetic datasets show that the metric scores decrease monotonically with an increasing number of injected errors, decline together with downstream ML performance, and scale linearly with the number of rows, which enables automated DQ monitoring based on the standards.
Comments14 pages, 7 figures, 5 tables. Submitted to EDBT 2027