发表机构
Fergana State Technical University(费尔干纳国立技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对概率分类器事后校准方法,通过条件分层稳健性分析,比较TEMP和ISO在四个条件下的表现,评估四个假设组,发现数据集中稳健性因条件和指标而异,未提及外部可迁移性。
AI 中文摘要
事后校准被广泛用于校正训练有素的分类器的概率估计,但大多数评估报告的是总体性能,而未测试在单个数据集中不同操作条件下该性能是否成立。我们进行了一项预先注册的条件分层稳健性分析,比较了温度缩放(TEMP)和等渗回归(ISO)在四个受控条件(C1 - C4)下的情况。评估了四个假设组:具有霍尔姆校正多重性控制的判别增量(H1)、布里尔得分差异(H2)、校准斜率结果(H3)以及最佳条件设置下的AUROC差异(H4)。结果表明,数据集中的稳健性取决于条件和指标,未涉及外部可迁移性。
英文摘要
Post-hoc calibration is widely adopted to correct probability estimates from trained classifiers, yet most evaluations report aggregate performance without testing whether that performance holds across distinct operating conditions within a single dataset. We present a pre-registered, condition-stratified robustness analysis comparing temperature scaling (TEMP) and isotonic regression (ISO) across four controlled conditions (C1--C4). Four hypothesis groups are evaluated: discrimination deltas with Holm-corrected multiplicity control (H1), Brier score differences (H2), calibration slope outcomes (H3), and AUROC differences under best-condition setups (H4). TEMP-minus-ISO discrimination deltas remain small across all conditions (-0.0155 to 0.0139), with Holm-adjusted p-values of 0.9895 everywhere. TEMP Brier differences are consistently negative (C1: -0.0002 through C4: -0.0074), while ISO shows sign reversals. TEMP calibration slopes stay closer to unity in every condition (range 0.7597--0.9493) than ISO slopes (0.1364--0.2726). AUROC differences shift from near zero in C1 (-0.0004) to positive in C4 (0.0264). These results establish that in-dataset robustness is condition-dependent and metric-specific. No claim of external transportability is made.
Comments6 pages, 5 figures