arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩展 GouDa:用于数据质量基准测试的带(和不带)错误的通用数据集生成

Extending GouDa: Generation of Universal Datasets with (and without) Errors for Data Quality Benchmarking

Valerie Restat, André Conrad, Kevin M. Kramer, Uta Störl

arXiv 2607.20165首次发表:更新:

AI 中文总结

研究围绕合成数据生成难题,提出GouDa数据生成器,可适配不同数据格式,能可控插入错误并生成基准事实,借助多种生成函数及属性值列表,生成涵盖多样用例的逼真数据,满足数据质量基准测试需求。

AI 中文摘要

合成数据在数据质量、数据清理和机器学习等领域极为重要,能助力分析真实数据不足、不可用或失真的用例。但生成合成数据面临挑战:数据要逼真且涵盖边缘情况,能插入可控错误并提供无错误版本,还要考虑多种数据格式。为此提出数据生成器GouDa,它适用于不同数据格式,能可控插入错误并生成基准事实,通过多种生成函数及添加属性值列表可生成涵盖多样用例的逼真数据。

英文摘要

Synthetic data is extremely important in areas such as data quality, data cleaning, and machine learning. It enables the analysis of use cases in which real data is insufficient, unavailable, or distorted. However, generating synthetic data also presents challenges: The data must be as realistic as possible, but at the same time cover edge cases. It must be possible to insert controlled errors, and at the same time, an error-free version of the data is usually required. Additionally, it is necessary to consider numerous data formats, such as tabular data, but also NoSQL data models. To this end, we present our data generator GouDa. GouDa precisely meets these requirements - it is suitable for different data formats, enables the controlled insertion of errors, and generates ground truth. A wide range of different generation functions and the option to add your own lists of possible attribute values allow the generation of realistic data that covers many different use cases.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑