发表机构
University of Colorado Denver; National Snow and Ice Data Center (NSIDC); University of Colorado Boulder(科罗拉多大学丹佛分校; 国家冰雪数据中心; 科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种保留多专家区间标签、分离标签内不精确性与专家间差异并验证预测不确定性成分的方法,在海冰浓度任务上相比硬标签降低MAE 31%。
AI 中文摘要
许多机器学习(ML)应用依赖于专家标签,而合格的专家可能对同一观测提供不同但合理的解释。这种专家标签间的差异可能反映真实的意见分歧或模糊性,而非标注错误。当个别专家额外报告区间而非精确值时,监督信号包含两种不同的标签不确定性来源:标签内不精确性和专家间差异。现有方法分别处理这些形式:多专家方法将标签折叠为共识,区间目标方法通常产生单一预测,而预测不确定性方法很少针对观测到的专家分歧验证其不确定性估计。为解决此问题,我们提出一种方法,保留个别专家区间,将标签内不精确性与专家间差异分离,并验证相应的预测不确定性成分。首先,异构标签词汇被协调为共同的概率标签空间,将编码差异与专家判断分离。其次,保留个别标签区间,并使用适当的Cramér距离目标训练的Beta分布混合模型进行建模,保留专家报告的独特标签。第三,我们将预测不确定性分解为成分内、成分间和模型不确定性,并评估这些成分是否分别对应于标签内不确定性、标签间不确定性和模型误差。由于这种对应关系不保证成立,我们引入分解匹配,将预测成分与其预期的标签侧来源对齐。在海冰浓度上,该模型相比硬标签将平均绝对误差(MAE)降低了31%,并优于聚合、区间分布和区间回归基线。
英文摘要
Many machine learning (ML) applications rely on expert labels, and qualified experts may provide different but plausible interpretations of the same observation. Such variation across expert labels may reflect genuine disagreement or ambiguity rather than annotation error. When individual experts additionally report intervals rather than exact values, the supervision contains two distinct sources of label uncertainty: within-label imprecision and between-expert variation. Existing methods treat these forms separately: multi-expert approaches collapse labels to a consensus, interval-target methods often yield a single prediction, and predictive-uncertainty methods rarely validate their uncertainty estimates against observed expert disagreement. To address this problem, we propose an approach that preserves individual expert intervals, separates within-label imprecision from between-expert variation, and validates the corresponding predictive uncertainty components. First, heterogeneous label vocabularies are harmonized into a common probabilistic label space, separating encoding differences from expert judgement. Second, individual label intervals are retained and modeled with a mixture of Beta distributions trained using a proper Cramér-distance objective, preserving distinct expert-reported labels. Third, we decompose predictive uncertainty into within-component, between-component, and model uncertainty, and evaluate whether these components correspond to within-label uncertainty, between-label uncertainty, and model error, respectively. Because this correspondence is not guaranteed, we introduce decomposition matching, which aligns the predictive components to their intended label-side sources. On sea-ice concentration the model reduces MAE by 31\% over hard labels and outperforms aggregation, interval-distribution, and interval-regression baselines.