arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越性能指标:Fazekas评分预测中标签模糊性的不确定性映射

Beyond Performance Metrics: Uncertainty Mapping of Label Ambiguity in Fazekas Score Prediction

Susanne Schmid, Johanna Ospel, Richard Frayne, Roberto Souza

arXiv 2609.17753首次发表:更新:

发表机构

University of Calgary; Schulich School of Engineering; Hotchkiss Brain Institute; Department of Radiology and Clinical Neurosciences; Calgary Image Processing and Analysis Centre; Seaman Family MR Research Centre(卡尔加里大学; 舒利希工程学院; 霍奇基斯脑研究所; 放射学与临床神经科学系; 卡尔加里图像处理与分析中心; 西曼家族磁共振研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出不确定性映射框架,将Fazekas评分预测模型的不确定性与特征表示位置关联,揭示类边界模糊区域和标签不一致,支持模型解释与数据集审查。

AI 中文摘要

用于训练医学图像分类模型的参考标签并不总是像它们看起来那样确定,这种不确定性对性能指标有影响。在本研究中,我们提出了一个框架来分析脑室周围Fazekas评分预测的模型性能,该框架超越了传统指标。Fazekas评分是一种序数视觉评定量表,用于评估白质高信号的严重程度,已知受评估者间变异性的影响。虽然最佳的Fazekas评分预测模型达到了马修斯相关系数(MCC)0.70,但性能在不同数据划分和损失函数之间有所变化,使得对模型能力的解释变得困难。我们的不确定性映射方法不是将模型预测的认知不确定性解释为孤立的标量值,而是将不确定性与其在学习到的特征表示中的位置相关联。这突出了类边界转换的区域,在这些区域中,病例显得更加模糊,误分类更可能发生。它还识别了潜在的标签不一致性,包括专家审查发现与原始参考Fazekas评分不一致的低不确定性误分类病例。因此,不确定性映射允许检查模型行为与类别分离和潜在的模型-标签不一致性的关系。损失函数的选择也影响了不确定性分布,一些模型显示出更清晰的类别分离和更局部化的模糊区域不确定性。这些发现表明,当参考标签受到模糊性/评估者间变异性影响时,Fazekas评分预测的不确定性映射可以支持模型解释和有针对性的数据集审查。

英文摘要

Reference labels used to train medical image classification models are not always as certain as they may appear, and this uncertainty has implications on performance metrics. In this study, we propose a framework to analyze model performance for periventricular Fazekas score prediction that goes beyond conventional metrics. The Fazekas score is an ordinal visual rating scale used to assess the severity of white matter hyperintensities and is known to be affected by inter-rater variability. While the best Fazekas score prediction model achieved a Matthews correlation coefficient (MCC) of 0.70, performance varied across data splits and loss functions, making interpretation of model capabilities difficult. Rather than interpreting epistemic uncertainty of a model's prediction as an isolated scalar value, our approach of uncertainty mapping relates uncertainty to its position within the learned feature representation. This highlights regions of class-boundary transitions where cases appear more ambiguous and misclassifications are more likely. It also identifies potential label disagreement, including low-uncertainty misclassified cases that expert review found to be inconsistent with the original reference Fazekas score. Therefore, uncertainty mapping allows model behaviour to be examined in relation to class separation and potential model-label disagreement. Loss function choice also influenced the uncertainty profile, with some models showing clearer class separation and more localized uncertainty in ambiguous regions than others. These findings suggest that uncertainty mapping for Fazekas score predictions can support model interpretation and targeted dataset review when reference labels are affected by ambiguity/ inter-rater variability.

Comments15 pages; 11 figures; 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑