arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22736cs.CV

评估领域差距:视频胶囊内镜领域转移下的模型选择不稳定性

Benchmarking the Domain Gap: Model Selection Instability Under Domain Shift in Video Capsule Endoscopy

Dan Hanson, Debesh Jha

首次发表
浏览论文内容

中文总结 AI 辅助

研究视频胶囊内镜领域转移下模型选择的不稳定性,通过在Kvasir - Capsule等数据集上微调预训练骨干网络并多目标评估,发现域内排名预测价值因目标而异,结论是模型选择应报告跨目标排名稳定性而非单数据集峰值性能。

中文摘要 AI 辅助

视频胶囊内镜(VCE)分类通常在单个数据集中进行评估,但临床应用需要在不同采集源、标注策略和患者群体中保持稳健性。我们使用Kvasir - Capsule、Capsule Vision 2024(CV2024)和Galar的共享标签子集来研究这一差距。在标准化协议下,我们在官方Kvasir - Capsule折叠上微调了一组通用领域预训练骨干网络,并在记录的共享标签决策空间内的两个非源目标上评估相同的检查点。我们发现域内排名的预测价值取决于目标:Kvasir - Capsule排名与Galar的一致性比与CV2024更高,而两个非源目标之间的一致性较弱。因此,最强的域内骨干网络在一个目标上领先,但在另一个目标上却处于中等水平,没有一个单一的评估目标能可靠地预测其他目标。第二个CV2024训练的配置集也重现了这种目标依赖的不稳定性。我们得出结论,胶囊内镜模型选择应报告跨目标排名稳定性,而不是单个数据集的峰值性能。

英文摘要

Video capsule endoscopy (VCE) classification is typically evaluated within a single dataset, yet clinical deployment demands robustness across acquisition sources, labeling policies, and patient populations. We examine this gap using Kvasir-Capsule, Capsule Vision 2024 (CV2024), and a shared-label subset of Galar. We fine-tune a suite of general-domain pretrained backbones on the official Kvasir-Capsule folds under a standardized protocol and evaluate the same checkpoints on two non-source targets within a documented shared-label decision space. We find that the predictive value of in-domain ranking is target-dependent: Kvasir-Capsule ranking aligns more closely with Galar than with CV2024, while the two non-source targets agree only weakly. Consequently, the strongest in-domain backbone leads on one target yet falls to mid-pack on the other, and no single evaluation target reliably predicts the others. A second CV2024-trained configuration set reproduces this target-dependent instability. We conclude that capsule endoscopy model selection should report cross-target ranking stability rather than peak single-dataset performance.

发表机构

  • University of South Dakota(南达科他大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑