发表机构
Faculty of Computer Science and Information Technology; Universiti Malaya(计算机科学与信息技术学院; 马来亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过72,000个分类问题,对比Jev和Laya模型在直接预测与通过类别重建预测间的一致性,发现显著不一致且影响各异,强调需联合评估准确率、校准和概率一致性。
AI 中文摘要
一个决策模型可以对每个问题给出总和为一的概率,但当同一决策被分解为更小的步骤时,却可能与自己不一致。我们在Jev和英语Laya检查点上研究了这种形式的概率一致性,每个系统在TREC、CLINC150和MASSIVE上使用了2,500个匹配示例。在72,000个分类问题中,我们比较了直接细粒度标签预测与宽泛类别概率以及通过这些类别重建的预测。两个系统都显示出显著的不一致:Jev的平均类别级总变差范围从0.219到0.349,Laya从0.424到0.689,在零表示完全一致的尺度上。后果差异显著。在CLINC150上,重建使Jev的准确率降低了22.9个百分点(配对95%自助法区间:[-24.9, -20.9]),而使Laya的准确率提高了21.3个百分点([18.0, 24.5])。相同的方向在所有三个数据集中保持一致,所有六个未调整的准确率变化区间均不包含零。准确率的提高也可能伴随着置信度可靠性的下降:在MASSIVE上,Laya获得了9.2个准确率点,而其期望校准误差从0.046上升到0.124。错误分析识别出宽泛类别错误和类别内混淆。这些发现表明,决策系统需要在应用程序使用的工作流程中对准确率、置信度校准和概率一致性进行联合评估。
英文摘要
A decision model can give probabilities that sum to one for every question yet disagree with itself when the same decision is broken into smaller steps. We study this form of probabilistic coherence in Jev and the English Laya checkpoint, using 2,500 matched examples per system across TREC, CLINC150, and MASSIVE. Across 72,000 classification questions, we compare direct fine-label predictions with broad-category probabilities and predictions reconstructed through those categories. Both systems show substantial disagreement: mean category-level total variation ranges from 0.219 to 0.349 for Jev and from 0.424 to 0.689 for Laya, on a scale where zero means exact agreement. The consequences differ sharply. On CLINC150, reconstruction reduces Jev's accuracy by 22.9 percentage points (paired 95% bootstrap interval: [-24.9, -20.9]) and improves Laya's by 21.3 points ([18.0, 24.5]). The same directions hold across all three datasets, with all six unadjusted accuracy-change intervals excluding zero. Improved accuracy can also accompany less reliable confidence: on MASSIVE, Laya gains 9.2 accuracy points while its expected calibration error rises from 0.046 to 0.124. Error analysis identifies both broad-category mistakes and within-category confusions. These findings show why decision systems need joint evaluation of accuracy, confidence calibration, and probability coherence in the workflow used by an application.
Comments20 pages, 4 figures. Code and experimental results: https://github.com/samanjoy2/system-one-coherence