重新思考准确性:一种基于加权误差的数据质量度量
Rethinking Accuracy: A Weighted Error-Based Metric for Data Quality
浏览论文内容
中文总结 AI 辅助
研究数据质量评估难题,提出TOMME这一基于加权误差的通用数据质量度量方法,能基于单一分数评估数据集质量,便于快速评估与自动化处理,还可依不同用例调整权重,可视为广义加权形式的准确性。
中文摘要 AI 辅助
实际数据常含错误,数据工程师需花大量时间创建数据清洗管道以确保数据质量。但比较不同管道结果并判断最佳者颇具难度,现有度量多针对特定用例,缺乏通用度量。本文提出TOMME,一种基于加权误差的通用数据质量度量初始方法。它能基于单一分数评估数据集质量,便于快速评估与自动化处理,还可通过不同权重精确适配特定用例,其名意为“测量误差的单一度量”,可视为广义加权形式的准确性。
英文摘要
Real data often contains errors, which is why data engineers spend a lot of time creating data cleaning pipelines to ensure the best possible data quality. However, it is often difficult to compare the results of different pipelines and decide which pipeline leads to the best results. There are many different metrics that are designed for different use cases, but they often take only a portion of the data into account. There is a lack of universally applicable metrics for measuring data quality that can be used in many different scenarios. That is why in this paper we are presenting TOMME - an initial approach to a universally applicable weighted error-based metric for data quality. This allows the data quality of a dataset to be assessed based on a single score. While a detailed data quality evaluation remains important, the use of a single score enables rapid assessment and automated processing, for example, for optimization algorithms. By using different weights, the score can also be precisely adjusted to the specific use case. That is why we named it TOMME, which stands for "The One Metric Measuring Errors". As the name suggests, it measures errors in the data. It can thus be considered a generalized, weighted form of accuracy.