发表机构
Lafayette College(拉斐特学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出先校准后融合的框架,利用合成序数参考空间对齐子集评分器,无需真值标签或共享标注空间,在三个基准上优于未校准平均,且降低计算成本。
AI 中文摘要
我们提出了一种先校准的框架,该框架无需访问真值标签或共享的标注空间即可生成监督分数。我们的框架在融合之前,使用一个合成的序数参考空间来对齐特定子集的评分器。该参考空间由代表潜在概念的有序校准特征构建,提供了一个公共尺度,使得原本不可比较的评分器输出可以在该尺度上对齐。由于我们的校准过程使用参考空间而非训练样本,因此它独立于训练集的经验分布。在三个基准数据集上,我们的框架始终优于未校准的平均方法,并且在评估指标上的主要指标点估计值高于最佳个体评分器。相对于基于样本的基线,性能因领域而异,在Ames Housing上的绝对差异低于0.02,在Breast Cancer Wisconsin和Wine Quality上低于0.01。经过Bonferroni校正后,Ames Housing上的所有三个比较以及Breast Cancer Wisconsin上的一个比较的差异仍然显著。此外,我们表明,每个特征使用较少的校准级别可以在大幅降低计算成本的情况下接近更高分辨率的结果。这些结果共同支持我们的框架作为在既无真值标签也无共享标注空间的情况下构建监督分数的一种可行方法。
英文摘要
We introduce a calibration-first framework that produces supervision scores without access to ground-truth labels or a shared annotation space. Our framework aligns subset-specific scorers using a synthetic ordinal reference space before fusion. This reference space is constructed from ordered calibration features that represent the latent concept, providing a common scale on which otherwise incomparable scorer outputs can be aligned. Because our calibration procedure uses the reference space rather than training samples, it is independent of the training set's empirical distribution. Across three benchmark datasets, our framework consistently outperforms uncalibrated averaging and achieves higher primary-metric point estimates on the evaluation metrics than the best individual scorer. Performance relative to sample-dependent baselines varies by domain, with absolute differences below 0.02 on Ames Housing and below 0.01 on Breast Cancer Wisconsin and Wine Quality. After Bonferroni correction, differences remain significant for all three comparisons on Ames Housing and one on Breast Cancer Wisconsin. Additionally, we show that using fewer calibration levels per feature can closely approximate higher-resolution results at substantially lower computational cost. Together, these results support our framework as a viable approach to construct supervision scores when neither ground-truth labels nor a shared annotation space is available.
Comments25 pages; 14 tables; 4 figures; 5 appendices; code available at https://github.com/Lafayette-EshbaughSilveyra-Group/calibration-first