LWCal:针对含噪校准标签的表格分类器的损失加权校准
LWCal: Loss-Weighted Calibration for Tabular Classifiers with Noisy Calibration Labels
- Brown University(布朗大学)
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对校准标签含噪的表格分类器,提出LWCal及其变体Gated-LWCal,通过损失加权和保守门控降低校准误差,实验证明优于现有方法。
AI中文摘要:
事后概率校准通常在一种乐观假设下进行评估:保留的校准标签是干净的。然而,在许多AI部署场景中,标签来自弱标注者、历史决策、启发式规则或远程监督,因此污染训练数据的相同标签噪声也会污染校准。我们研究了这一被忽视的表格分类器失效模式,并提出了LWCal,一种仅使用CPU的事后校准器,它降低那些噪声标签与基础模型保留概率相矛盾的校准样本的权重。LWCal不需要干净的验证标签、不需要噪声率估计,也不需要重新训练基础分类器。第二种变体Gated-LWCal增加了一个保守的不一致门控,当校准分割显得极度不一致时,它会回退到原始分数。在九个本地二元表格任务、六个随机种子、对称和非对称标签损坏以及三种基于树的基学习器上,LWCal获得了最低的平均校准误差,而Gated-LWCal获得了最佳的平均适当分数权衡。在主要的随机森林研究中,覆盖432个噪声单元,Gated-LWCal将期望校准误差从0.188降至0.122,负对数似然从0.438降至0.396,相对于原始分类器。Gated-LWCal与原始、Platt、等渗和beta校准的配对自助置信区间在ECE、Brier分数和NLL上均不包含零。该工件包含所有脚本、结果表、图表和编译后的论文。
英文摘要:
Post-hoc probability calibration is usually evaluated under an optimistic assumption: the held-out calibration labels are clean. In many AI deployment settings, however, labels come from weak annotators, historical decisions, heuristics, or distant supervision, so the same label noise that corrupts training also corrupts calibration. We study this overlooked failure mode for tabular classifiers and propose LWCal, a CPU-only post-hoc calibrator that down-weights calibration examples whose noisy labels are contradicted by the base model's held-out probability. LWCal requires no clean validation labels, no noise-rate estimate, and no retraining of the base classifier. A second variant, Gated-LWCal, adds a conservative disagreement gate that backs off toward the raw score when the calibration split appears extremely inconsistent. On nine local binary tabular tasks, six random seeds, symmetric and asymmetric label corruption, and three tree-based base learners, LWCal obtains the lowest average calibration error while Gated-LWCal obtains the best average proper-score tradeoff. In the main random-forest study over 432 noisy cells, Gated-LWCal reduces expected calibration error from 0.188 to 0.122 and negative log likelihood from 0.438 to 0.396 relative to the raw classifier. Paired bootstrap intervals for Gated-LWCal versus raw, Platt, isotonic, and beta calibration exclude zero on ECE, Brier score, and NLL. The artifact contains all scripts, result tables, figures, and the compiled paper.