AI 中文总结
研究围绕合成数据生成难题,提出GouDa数据生成器,可适配不同数据格式,能可控插入错误并生成基准事实,借助多种生成函数及属性值列表,生成涵盖多样用例的逼真数据,满足数据质量基准测试需求。
AI 中文摘要
合成数据在数据质量、数据清理和机器学习等领域极为重要,能助力分析真实数据不足、不可用或失真的用例。但生成合成数据面临挑战:数据要逼真且涵盖边缘情况,能插入可控错误并提供无错误版本,还要考虑多种数据格式。为此提出数据生成器GouDa,它适用于不同数据格式,能可控插入错误并生成基准事实,通过多种生成函数及添加属性值列表可生成涵盖多样用例的逼真数据。
英文摘要
Synthetic data is extremely important in areas such as data quality, data cleaning, and machine learning. It enables the analysis of use cases in which real data is insufficient, unavailable, or distorted. However, generating synthetic data also presents challenges: The data must be as realistic as possible, but at the same time cover edge cases. It must be possible to insert controlled errors, and at the same time, an error-free version of the data is usually required. Additionally, it is necessary to consider numerous data formats, such as tabular data, but also NoSQL data models. To this end, we present our data generator GouDa. GouDa precisely meets these requirements - it is suitable for different data formats, enables the controlled insertion of errors, and generates ground truth. A wide range of different generation functions and the option to add your own lists of possible attribute values allow the generation of realistic data that covers many different use cases.