AI 中文总结
本研究在三种噪声场景下对比不确定性采样与随机采样,发现其标签效率受数据集、预算等因素影响,未找到结构化错误定位带来普遍额外惩罚的证据。
AI 中文摘要
主动学习可通过选择信息丰富的样本来降低标注成本,但最不确定的样本也可能最难被正确标注。本研究测试不确定性采样失效的原因是其获取了更多被损坏的标签,还是错误集中在困难区域造成的危害更大。在干净标签、随机分类噪声(RCN)以及与难度相关的有界噪声三种场景下,在三个公开的二元表格数据集上,将基于边际的不确定性采样与随机采样进行对比。实验设计采用100对配对随机种子、0至0.30范围内的9种期望噪声率、20至120的标注预算,以及在每个预算下通过交叉验证重新选择正则化的逻辑回归模型。一个与曝光匹配的RCN对照组对齐了最终获取的平均损坏程度,而干净标签的扩展版本达到了400的预算。在干净标签下,不确定性采样在所有数据集上将学习曲线下的归一化平衡准确率面积提高了1.09至1.77个百分点。在威斯康星乳腺癌数据集的8种噪声率中有6种,与难度相关的噪声比RCN更大程度地降低了这一优势,但在纸币鉴别或MAGIC伽马望远镜数据集的所有测试噪声率下均未出现该情况。曝光匹配分析未发现结构化错误定位带来普遍额外惩罚的修正证据。在干净的MAGIC数据上,不确定性采样提高了平衡准确率,但在固定假阳性率下降低了平均精度和真阳性率。因此,不确定性采样具有标签效率,但其表现出的鲁棒性取决于数据集、预算、噪声结构和评估指标。
英文摘要
Active learning can reduce labeling cost by selecting informative examples, but the most uncertain examples may also be the hardest to label correctly. This study tests whether uncertainty sampling fails because it acquires more corrupted labels or because errors concentrated in difficult regions are especially harmful. Margin-based uncertainty sampling is compared with random sampling under clean labels, random classification noise (RCN), and bounded difficulty-dependent noise on three public binary tabular datasets. The design uses 100 paired seeds, nine expected noise rates from 0 to 0.30, annotation budgets from 20 to 120, and logistic regression with regularization re-selected by cross-validation at every budget. An exposure-matched RCN control aligns mean final acquired corruption, while a clean-label extension reaches budget 400. Under clean labels, uncertainty sampling improved normalized balanced-accuracy area under the learning curve by 1.09 to 1.77 percentage points on all datasets. Difficulty-dependent noise reduced this advantage more than RCN at six of eight rates on Breast Cancer Wisconsin, but at no tested rate on Banknote Authentication or MAGIC Gamma Telescope. Exposure-matched analyses found no corrected evidence for a universal additional penalty from structured error location. On clean MAGIC data, uncertainty sampling improved balanced accuracy while reducing average precision and true-positive rate at fixed false-positive rates. Thus, uncertainty sampling was label-efficient, but its apparent robustness depended on dataset, budget, noise structure, and evaluation metric.
Comments14 pages, 5 figures, 4 tables. Code and reproducibility materials: https://github.com/dev-juy/hard-cases-bad-labels