arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06053cs.LG

基于大语言模型的数据质量规则生成

Data Quality Rule Generation with LLMs

Anna-Christina Glock, Thomas Hütter, Johannes Fürnkranz, Wolfram Wöß, Christine Dominka-Kiss, Lisa Ehrlinger

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出LeDQeR,一种基于大语言模型的数据质量规则自动生成方法,通过生成-过滤框架生成候选规则并应用四种过滤技术,确保规则有效且紧凑。

中文摘要 AI 辅助

数据验证,例如客户数据和员工数据的验证,在许多组织中是一项重要任务。数据中的错误可能带来严重后果。例如,患者记录中错误的药物单位可能导致危及生命的用药错误,地址中缺失街道号码可能导致投递失败。公司通常采用基于规则的企业级数据质量(DQ)工具,这些工具允许领域专家指定规则以随时间验证数据。虽然基于规则的DQ工具计算效率高且提供可解释的报告,但手动维护一套全面的规则集具有挑战性,因为领域专家常常忽略关键规则,尤其是在复杂领域和大数据量情况下。因此,在实践中弥合这些差距仍然是一个开放问题。在本文中,我们应对自动DQ规则生成的挑战。为此,我们形式化了一个通用的生成-过滤框架,并引入了LeDQeR,一种基于大语言模型的DQ规则生成方法。首先,大语言模型(LLM)根据观察到的脏数据元组,针对给定的基于规则的DQ工具语法生成候选规则。其次,我们应用四种过滤技术,以确保生成规则的(i)可执行性,(ii)正确性和(iii)泛化性,并避免(iv)冗余。广泛的实验评估表明,LeDQeR能够为各种数据集和错误类型生成有效且紧凑的规则集。

英文摘要

The validation of data, such as customer and employee data, is an important task in many organizations. Errors in data can have severe consequences. For example, a wrong drug unit in a patient record can lead to life-threatening medication errors, and a missing street number in an address to failed deliveries. Companies often employ rule-based enterprise data quality (DQ) tools, which allow domain experts to specify rules to validate the data over time. While rule-based DQ tools are computationally efficient and provide explainable reports, maintaining a comprehensive rule set manually is challenging, as domain experts often overlook essential rules, especially in complex domains and large data volumes. Hence, closing these gaps remains an open problem in practice. In this paper, we address the challenge of automated DQ rule generation. For this, we formalize a generalizable generate-filter framework and introduce LeDQeR, an LLM-based DQ rule generation approach. First, a large language model (LLM) generates candidates rules from an observed dirty data tuple for a given rule-based DQ tool syntax. Second, we apply four filter techniques that ensure the (i) executability, (ii) correctness, and (iii) generalizability, and avoid (iv) redundancy of the generated rules. An extensive experimental evaluation suggests that LeDQeR is able to produce effective and compact rule sets for various datasets and error types.

发表机构

  • Software Competence Center(软件能力中心)
  • Johannes Kepler University(约翰内斯·开普勒大学)
  • Austrian Post(奥地利邮政)
  • Hasso Plattner Institute, University of Potsdam(哈索·普拉特纳研究所,波茨坦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑