arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33457cs.LG

一个自由的旋钮:在基于阈值的评估中解耦校准与预测技能

A Free Knob: Decoupling Calibration and Predictive Skill in Threshold-Based Evaluation

Md Tanveer Hossain Munim, Bijoy Ahmed Saiem, Al-Amin Sany, Tanzima Hashem

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出FreeKnob审计,通过单调校准解耦阈值评估中的校准与预测技能,证明CSI混淆源于失准,并建议报告频率偏差和对称审计。

中文摘要 AI 辅助

许多密集预测基准通过将预测和目标在空间块上池化、对每个块进行阈值化并对列联表打分来评估罕见事件。在固定的罕见工作点上,最大池化临界成功指数(CSI)将空间判别力与幅度校准混为一谈:尖锐的观测使许多块超过阈值,而平方误差回归产生的衰减预测使相同的块低于阈值。我们将经典的单调校准重新用作对称审计:在留出数据上拟合的后处理变换,并分别应用于每个系统。该变换无法逆转像素排序,因此它复现的任何对比都不能确立改进的空间排名。在SEVIR上,同一架构的两个已发布检查点在控制前极端阈值CSI相差-29.5%,在控制后相差+5.3%。在6个系统的450个成对对比中,池化频率偏差差异的变化与控制下CSI对比移动的距离相关(r = +0.796),51个对比反转了符号。在CasCast发布的极端事件工作点上,级联-骨干CSI差距从0.1601降至0.0339,减少了78.8%;剩余差距保持为正。当变换在测试期之前的窗口上拟合时,该效应持续存在,校准还揭示了被更好校准的基线所隐藏的优势。在地球静止红外图像上,相对增益随着事件变得更罕见而增长;人群计数在补丁求和池化下重现了偏差-增益关系;语义分割中频率偏差已接近1,平均变化很小。因此,这种混淆需要固定工作点和使输出在该点失准的训练机制。我们建议在罕见事件池化和阈值分数旁边报告池化频率偏差和对称的留出FreeKnob审计。

英文摘要

Many dense-prediction benchmarks evaluate rare events by pooling prediction and target over spatial blocks, thresholding each, and scoring the contingency table. At a fixed rare operating point, the max-pooled Critical Success Index (CSI) confounds spatial discrimination with amplitude calibration: sharp observations promote many blocks above threshold, while attenuated predictions from squared-error regression leave the same blocks below it. We repurpose classical monotone calibration as a symmetric audit: a post-hoc transform fitted on held-out data and applied separately to each system. The transform cannot reverse pixel ordering, so any contrast it reproduces cannot establish improved spatial ranking. On SEVIR, two released checkpoints of one architecture differ by -29.5% in extreme-threshold CSI before the control and by +5.3% after it. Across 450 pairwise contrasts among 6 systems, the difference in pooled frequency-bias deviation is associated with how far the CSI contrast moves under the control (r = +0.796), and 51 contrasts reverse sign. At CasCast's published extreme-event operating point, the cascade-over-backbone CSI gap falls from 0.1601 to 0.0339, a 78.8% reduction; the remaining gap stays positive. The effect persists when the transform is fitted on a window before the test period, and calibration also reveals advantages hidden by a better-calibrated baseline. On geostationary infrared imagery the relative gain grows as events become rarer, crowd counting reproduces the bias-gain relationship under patch-sum pooling, and semantic segmentation, where frequency bias is already near one, shows little average change. The confound therefore requires both a fixed operating point and a training regime that leaves the output miscalibrated there. We recommend reporting pooled frequency bias and a symmetric held-out FreeKnob Audit alongside rare-event pool-and-threshold scores.

发表机构

  • Regional Integrated Multi-Hazard Early Warning System (RIMES)(区域综合多灾种早期预警系统)
  • Bangladesh University of Engineering and Technology (BUET)(孟加拉国工程技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑