arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13711eess.IVcs.CVcs.LG

TRUE-Colon:揭示实时息肉检测中一致的迁移不对称性

TRUE-Colon: Exposing a Consistent Transfer Asymmetry in Real-Time Polyp Detection

Sebastian Doerrich, Andreas Franz Schwab, Francesco Di Salvo, Shyam Nandan Rai, Hanh Huyen My Nguyen, Christian Ledig

首次发表
浏览论文内容

中文总结 AI 辅助

本研究建立TRUE-Colon基准协议,发现实时息肉检测模型存在迁移不对称性,经完整流程训练的模型表现更优,建议CADe训练及基准测试转向完整流程数据。

中文摘要 AI 辅助

用于结肠镜检查的计算机辅助检测(CADe)系统有望降低临床漏检率,但可靠的实际部署仍难以实现。这种转化差距部分源于模型开发中的结构性缺陷:依赖精选数据集,这类数据集未能充分代表常规检查中常见的长阴性段和操作相关伪影。严格在这些以病灶为中心的基准上训练和评估架构会造成成功的假象,因为此类基准无法捕捉临床关键指标。为揭示这一差距,我们建立了TRUE-Colon,这是一种标准化基准测试协议,可在定位精度之外测量关键部署特性,并在精选基准(SUN、PICCOLO)和60个未编辑的完整操作流程(REAL-Colon)上评估四种实时架构(Faster R-CNN、YOLOv8、YOLOv11、RT-DETR)。我们观察到一致的迁移不对称性:严格在精选片段上训练的模型在完整流程评估时会出现严重的性能崩溃,而经流程训练的模型在REAL-Colon上对非息肉内容的拒斥能力显著提升,且在精选基准上基本保持精度。除可迁移性外,我们还发现Transformer检测器达到最强灵敏度,且检测最早、最持久,而卷积检测器在更高吞吐量上保持竞争力。综合来看,这些结果表明,可部署CADe的训练和基准测试应从精选的以病灶为中心的片段转向完整流程数据和与部署相关的操作点。源代码可在该https URL获取。

英文摘要

Computer-aided detection (CADe) systems for colonoscopy promise to reduce clinical miss rates, yet reliable real-world deployment remains elusive. This translational gap stems in part from a structural flaw in model development: the reliance on curated datasets that under-represent the long negative stretches and procedure-related artifacts characteristic of routine examinations. Training and evaluating architectures strictly on these lesion-centric benchmarks creates an illusion of success, since such benchmarks cannot capture clinically crucial metrics. To expose this gap, we establish TRUE-Colon, a standardized benchmarking protocol that measures key deployment characteristics alongside localization accuracy, and evaluate four real-time architectures (Faster R-CNN, YOLOv8, YOLOv11, RT-DETR) across curated benchmarks (SUN, PICCOLO) and 60 unedited, full-length procedures (REAL-Colon). We observe a consistent transfer asymmetry: models trained strictly on curated clips suffer a severe performance collapse when evaluated on full procedures, whereas procedure-trained models substantially improve rejection of non-polyp content on REAL-Colon, and largely retain their accuracy on curated benchmarks. Beyond transferability, we find that the Transformer detector attains the strongest sensitivity and the earliest, most persistent detections, while the convolutional detectors stay competitive at a higher throughput. Together, these results indicate that both training and benchmarking for deployable CADe should shift from curated, lesion-centric clips toward full-procedure data and deployment-relevant operating points. Source code is available at https://github.com/sdoerrich97/true-colon.

发表机构

  • University of Bamberg(班贝格大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑