计算病理学中伪影检测的可靠基准:可复现性与不确定性分析
Reliable Benchmarking of Artifact Detection in Computational Pathology: A Reproducibility and Uncertainty Analysis
浏览论文内容
中文总结 AI 辅助
本研究针对计算病理学伪影检测基准的缺陷,提出含四类变异检验的可靠性协议,经评估发现原基准的核心机制可复现但对比声明不可靠,小队列基准结论可信度低。
中文摘要 AI 辅助
背景与目的:质量控制是全切片图像分析的前提,但用于对比质量控制方法的基准具有四个特性,导致其报告的差异难以解读:独立切片数量少、注释集中在少数切片中、无闭式标准误的合并比率指标、单一继承的训练/测试划分。我们为此类基准提出了一种可靠性协议。方法:该协议量化了四类变异来源——测试集采样、训练随机性、划分构成以及未记录的预处理;只有当声明通过所有四类检验时,才具备可报告性;其中三类检验仅需数分钟的计算资源。我们将其应用于对已发表的基于扩散的伪影检测器的独立重建,在原始24切片划分上进行评估,并与一个有监督基线模型对比。结果:该方法的核心机制可复现:辅助对比项将合并F1值从0.673提升至0.688,且在第二个随机种子下可重复(提升0.0156,p=0.031;提升0.0190,p=0.005),不过其作用对象是笔痕,而非用于激发该方法的伪影类型。其对比声明不可复现:设计变体之间的差异,以及与有监督基线模型的差异,均处于评估的不确定性范围内。24张切片中有4张承载了70%的已评分注释像素,有效样本量为6.2,而继承的划分处于第7百分位。一项未报告的组织限制步骤排除了41.4%的离焦注释,仅排除2.6%的气泡注释;此类门控机制在结构上与模糊效应相混淆。结论:小队列基准能支持的结论远弱于当前报告所暗示的程度。这四项检验成本低廉,足以伴随任何对此类资源的评估,可将可复现效应与评估无法解析的差异区分开来。
英文摘要
Background and Objective: Quality control is a prerequisite for whole-slide image analysis, yet the benchmarks on which quality-control methods are compared share four properties that make their reported differences hard to interpret: few independent slides, annotation concentrated in a minority of them, pooled ratio metrics with no closed-form standard error, and a single inherited train/test partition. We propose a reliability protocol for such benchmarks. Methods: The protocol quantifies four sources of variability - test-set sampling, training stochasticity, partition composition, and undocumented preprocessing - a claim is reportable only if it survives all four; three of the four cost minutes of compute. We apply it to an independent reconstruction of a published diffusion-based artifact detector, evaluated on the original 24-slide partition and against a supervised baseline. Results: The method's central mechanism reproduces: the auxiliary contrastive term improves pooled F1 from 0.673 to 0.688 and replicates under a second seed (+0.0156, p = 0.031; +0.0190, p = 0.005), although it acts on pen marking rather than the artifact types cited to motivate it. Its comparative claims do not: differences between design variants, and against the supervised baseline, fall inside the uncertainty of the evaluation. Four of 24 slides carry 70% of scored annotated pixels, giving an effective sample size of 6.2, and the inherited partition sits at the 7th percentile. An unreported tissue-restriction step excludes 41.4% of out-of-focus annotation against 2.6% of air bubble; such a gate is confounded with blur by construction. Conclusions: Small-cohort benchmarks support far weaker conclusions than current reporting implies. The four checks are cheap enough to accompany any evaluation on such a resource and separate reproducible effects from differences the evaluation cannot resolve.
发表机构
- University of Piraeus(比雷埃夫斯大学)
机构由 AI 辅助整理,请以论文原文为准。