arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CURED:创建、理解和修复错误演示器

CURED: Creating, Understanding, and Repairing Errors Demonstrator

Nicholas Chandler, Sebastian Jäger, Philipp Jung, Felix Bießmann

arXiv 2607.20140首次发表:更新:

发表机构

Berliner Hochschule für Technik; Einstein Center Digital Future(柏林工业大学; 爱因斯坦数字未来中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究表格数据错误检测与清理,结合基于机器学习的数据清理和错误模型工作,创建统一演示器,用户能上传数据、引入错误并清理,弥合了错误模型和数据清理算法理论与实践的差距。

AI 中文摘要

检测和清理表格数据中的错误是数据密集型软件应用程序的前提。机器学习(ML)和数据库管理系统(DBMS)交叉领域的最新研究突出了统计学习算法在错误检测和清理方面的潜力。本文在一个统一的演示器中结合了我们近期基于ML的数据清理和错误模型的工作。该Web应用程序允许用户上传表格数据,用实际的数据相关错误干扰数据,并使用现代ML方法清理和理解数据中的错误机制。我们的演示器有助于弥合表格数据错误模型和数据清理算法在理论进展与直观实践见解之间的差距。演示器可通过此https链接获取。

英文摘要

Detecting and cleaning errors in tabular data is a prerequisite for data intense software applications. Recent research at the intersection of Machine Learning (ML) and Database Management Systems (DBMS) highlights the potential of statistical learning algorithms for error detection and cleaning. This paper combines our recent work on ML-based data cleaning and error models in a unified demonstrator. The web application allows users to upload tabular data, perturb the data with realistic data dependent errors and use modern ML methods to clean and understand error mechanisms in data. Our demonstrator helps to bridge the gap between theoretical advancements and intuitive practical insights in the context of error models and data cleaning algorithms for tabular data. The demonstrator is available at https://cured.demo.calgo-lab.de/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑