AI 中文总结
研究针对AI辅助结肠镜检查中仅置信度不能全面评估预测的问题,提出EndoExplain审计框架,整合多种要素。通过实验展示了分类器和分割器的性能,还验证了归因方法等,该框架能分离多种信号,为临床审查提供支持,是回顾性研究原型。
AI 中文摘要
人工智能辅助结肠镜检查系统通常报告帧级置信度分数,但仅置信度并不能表明预测在空间上是否合理、在时间上是否持续或在图像质量下降时是否可靠。我们提出了EndoExplain,一种用于内镜人工智能审查的轻量级且可重复的审计框架。该框架整合了分类置信度、病变分割、CAM风格的视觉归因、归因掩码对齐、帧质量指标和时间事件总结,旨在分离在计算机辅助检测管道中经常混淆的信号。在HyperKvasir上,所选的EfficientNet - B0分类器在十个内镜类别上的测试准确率达到0.9280,息肉家族与其他视图的ROC - AUC为0.9969。所选的U - Net++ EfficientNet - B1分割器在分割测试集上的Dice为0.9318,IoU为0.8826。严格的前20%多方法归因审计表明,归因方法强烈改变解释掩码一致性:Eigen - CAM重叠最强,而Grad - CAM++对齐较弱。对ETIS - LaribPolypDB和CVC - ClinicDB的冻结模型外部合理性检查保留了这种归因方法排名,同时显示了数据集相关的重叠。对60个HyperKvasir视频进行的人工审查的剪辑级时间基准测试,使用重叠≥1秒的一对一事件匹配,在预先指定的阈值0.85下达到事件F1 0.8081。临床医生告知的外部合理性审查支持了分离置信度、定位、归因、质量元数据和时间上下文的临床可读性。由此产生的驾驶舱式审查层将这些信号呈现为不同的可审计输出。这是一个回顾性研究原型,而不是经过临床验证的医疗设备。
英文摘要
Background and objective: A high classifier score and a plausible class-activation map (CAM) are often presented together, although neither establishes that the other is reliable. We introduce endoExplain as a reproducible protocol for auditing score-localisation discordance rather than as a new detector or explanation algorithm. Methods: Content hashing separated HyperKvasir development images from 1,000 masked images before training. EfficientNet-B0, ResNet-34 and ConvNeXt-Tiny were trained with three seeds each. Scores were temperature scaled using validation data only. Grad-CAM, Grad-CAM++, XGrad-CAM, HiResCAM and Eigen-CAM were evaluated on identical image-mask pairs, alongside random and centre baselines. Outcomes combined peak localisation, overlap, a top-20% deletion response, score-threshold sensitivity and adjustment for lesion size and centrality. The selected checkpoint was transferred without retraining to three external mask cohorts. Results: Temperature scaling reduced test expected calibration error from 0.0167 to 0.0115. Among 172 reserved images with scaled score at least 0.90, peak-outside-lesion rates ranged from 4.1% to 62.2% across CAMs. Method dependence remained evident across architectures and seeds, although method rankings were not universal. Spatial alignment and deletion response were not interchangeable. A random-deletion control also produced positive logit drops, limiting specificity claims based on deletion alone. External positive-mask results were dataset dependent. A source-category audit also exposed that 149/155 test positives were dyed-lifted polyps, materially bounding classifier claims. Conclusions: endoExplain makes calibration, spatial agreement, perturbation response and transfer separately inspectable. The results caution against using a score or visually persuasive CAM as evidence of lesion localisation or model reasoning.
Comments18 pages, 5 figures, 5 tables. Substantially revised version with a new title and scope; added calibration analysis, paired multi-CAM evaluation, random-deletion controls, architecture and seed robustness analyses, and frozen external spatial transfer