arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06511cs.LG

揭示移除预算混杂:一种用于自适应数据清洗的匹配操作点评估框架

Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning

Wei-Hsiang Chen, Pin-Hsuan Yu, Chen-Hsuan Fang, Jung-Hua Wang

AI总结:

该研究针对自适应数据清洗中存在的移除预算混杂问题,提出匹配操作点评估框架,经 CIFAR-10、ImageNet-100 实验验证,可消除虚假性能提升,确保结果反映真实损坏区分能力。

AI中文摘要:

自适应数据清洗方法用数据驱动的划分取代手动过滤阈值。然而,划分粒度(按估计的损坏风险对样本进行分割的组数)的变化会隐性改变决策边界,并改变移除样本的总数。这会产生一种称为移除预算混杂的偏差,其中精度或假阳性率等指标的表观提升反映的是更小的移除预算,而非更优的损坏区分能力。为解决这种评估偏差,我们引入了一种感知操作点的评估框架,该框架使用匹配预算和匹配召回率控制,以及与阈值无关的指标(AUROC 和 AUPRC)来评估方法。我们在一个多线索自适应清洗器的重新设计上测试了该框架,该重新设计包含一个加权的学习难度线索、一个辅助欧氏距离线索,以及旨在分离干净但难以分类样本的更细划分粒度。虽然朴素评估(在各自诱导的操作点评估配置)表明该重新设计有显著的性能提升,但当操作点均等时,这些提升消失了。假阳性分解显示,干净但难以分类的样本主要在低损坏率下驱动误差计数,在中等损坏率下成为阈值依赖,在严重损坏下贡献可忽略不计。在 CIFAR-10 和 ImageNet-100 上的实验表明,朴素评估中观察到的大多数性能差异在低至中等损坏率下,当操作点匹配时会缩小或消失。真正的排名优势仅保留在特定的低流行度设置和严重损坏下的高召回区域。这些发现强调,自适应清洗方法必须在匹配的操作点上进行基准测试,以确保性能提升反映真实的损坏区分能力。

英文摘要:

Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly shift the decision boundary and alter the overall number of removed samples. This creates a bias known as removal-budget confounding, where apparent gains in metrics like precision or false-positive rate reflect a smaller removal budget rather than superior corruption discrimination. To address this evaluation bias, we introduce an operating-point-aware evaluation framework that evaluates methods using matched-budget and matched-recall controls alongside threshold-independent metrics (AUROC and AUPRC). We test this framework on a multi-cue adaptive cleaner redesign featuring a reweighted learning-difficulty cue, an auxiliary Euclidean-distance cue, and increased partition granularity intended to isolate clean-but-difficult samples. While naive evaluations (assessing configurations at their own induced operating points) suggest substantial performance improvements for the redesign, these gains disappear once operating points are equalized. False-positive decomposition reveals that clean-but-difficult samples primarily drive error counts at low corruption rates, become threshold-dependent at moderate corruption, and contribute negligibly under severe corruption. Experiments on CIFAR-10 and ImageNet-100 demonstrate that most performance differences observed in naive evaluation shrink or vanish at low-to-moderate corruption when operating points are matched. True ranking advantages only remain in specific low-prevalence settings and in high-recall regions under severe corruption. These findings highlight that adaptive cleaning methods must be benchmarked at matched operating points to ensure performance gains reflect genuine corruption discrimination.

↑