AI 中文总结
针对表格数据中的噪声标签问题,提出一种数据为中心的元模型自适应集成选择方法,通过预测检测器权重来诊断数据质量并选择清洗策略,在受控噪声基准上达到与Confident Learning相当的性能。
AI 中文摘要
表格数据集中不正确或损坏的标签会显著降低监督学习的性能,尤其是当错误标注较为隐蔽且不易仅从特征空间检测到时。在自动化或AI增强的数据科学工作流中,对这种标签噪声的稳健检测对于构建可靠模型至关重要。我们提出了一种面向AI数据科学系统的数据为中心推理模块,该模块自动诊断数据集质量并选择合适的清洗策略。给定一个数据集,一个元模型在一组多样化的检测器上预测权重,这些检测器包括基于置信度的方法、基于邻域的方法和基于分布的方法。在具有受控噪声的基准数据集上,我们的方法平均达到了与Confident Learning基线相当的性能,但存在依赖数据集的收益和损失,特别是在异质性较强的场景中。我们进一步表明,检测器的有效性系统地与数据集属性相关联。这些结果证明了描述符驱动的、以数据为中心的集成作为AI辅助数据科学流程中用于稳健数据集评估和模型可靠性的组件的价值。
英文摘要
Incorrect or corrupted labels in tabular datasets can significantly degrade supervised learning performance, particularly when mislabeling is subtle and not easily detectable from feature space alone. In the context of automated or AI-augmented data science workflows, robust detection of such label noise is critical for building reliable models. We propose a data-centric reasoning module for AI data science systems that automatically diagnoses dataset quality and selects appropriate cleaning strategies. Given a dataset, a meta-model predicts weights over a diverse set of detectors, including confidence-based, neighborhood-based, and distributional methods. Across benchmark datasets with controlled noise, our approach achieves performance comparable to a Confident Learning baseline on average, with dataset-dependent gains and losses, particularly in heterogeneous regimes. We further show that detector effectiveness is systematically linked to dataset properties. These results demonstrate the value of descriptor-driven, data-centric ensembling as a component of AI-assisted data-science pipelines for robust dataset assessment and model reliability.