发表机构
Indian Institute of Technology Roorkee(印度理工学院鲁尔基分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于裁判的质量估计(RBQE),利用主模型与独立裁判模型的一致性作为部署时可靠性信号,在外部基准上显著提升息肉分割的可靠性评估与选择性预测。
AI 中文摘要
在实时结肠镜检查中,推理时无法获得真实标注,因此息肉分割模型可能无声地失败。我们提出了基于裁判的质量估计(RBQE),这是一个无参考框架,通过测量主分割模型与独立训练的裁判模型在同一图像上的一致性来评估可靠性。RBQE在从四个公共数据集构建的标准化1,223张图像外部基准上进行了评估,使用了四种裁判配置,以分离两个设计轴:裁判独立性和架构多样性。使用通用的Agreement Dice描述符,仅随机初始化不同但架构相同的裁判模型已能产生有用的可靠性信号(ROC-AUC = 0.923),表明仅独立训练就足够了。跨架构裁判模型进一步提升了性能:SegFormer-B0取得了最强性能(ROC-AUC = 0.960),显著优于同架构对照和UNet++,并在相同协议下比代表性测试时增强基线高出0.055 ROC-AUC,而提示耦合的MedSAM裁判尽管具有最大架构多样性,表现却较差。由于空掩码一致性可平凡分离,我们还报告了排除此类情况的受限评估:ROC-AUC降至0.876(SegFormer-B0,1,046张图像)和0.783(同架构对照,975张图像),但在此相同子集上,RBQE相对于两个基线的优势反而扩大。此外,随着低一致性案例被逐步拒绝,RBQE提高了保留预测的平均Dice,支持选择性预测,并且在推理时仅需额外一次确定性裁判前向传播。因此,我们的研究支持跨模型一致性作为自动息肉分割的实用、可解释的可靠性框架。
英文摘要
In real-time colonoscopy, ground-truth annotations are unavailable at inference, so polyp segmentation models can fail silently. We propose Referee-Based Quality Estimation (RBQE), a reference-free framework measuring agreement between a primary segmentation model and an independently trained referee on the same image. RBQE is evaluated on a standardized 1,223-image external benchmark drawn from four public datasets, using four referee configurations chosen to separate two design axes: referee independence and architectural diversity. Using a common Agreement Dice descriptor, a same-architecture referee differing from the primary model only in random initialization already yields a useful reliability signal (ROC-AUC = 0.923), showing that independent training alone is sufficient. Cross-architecture referees improve further: SegFormer-B0 achieves the strongest performance (ROC-AUC = 0.960), significantly outperforming the same-architecture control and UNet++, and exceeding a representative Test-Time Augmentation baseline by 0.055 ROC-AUC under an identical protocol, whereas a prompt-coupled MedSAM referee underperforms despite maximal architectural diversity. Because empty-mask agreement is trivially separable, we also report a restricted evaluation excluding such cases: ROC-AUC falls to 0.876 (SegFormer-B0, 1,046 images) and 0.783 (same-architecture control, 975 images), yet RBQE's margin over both baselines widens on this identical subset. RBQE additionally increases the mean Dice of retained predictions as low-agreement cases are progressively rejected, supporting selective prediction, and requires only one additional deterministic referee forward pass at inference. Our study therefore supports cross-model agreement as a practical, interpretable reliability framework for automated polyp segmentation.