CasCVS-Net:用于安全关键视图评估的分阶段多任务级联网络
CasCVS-Net: A Staged Multi-Task Cascade for Critical View of Safety Assessment
浏览论文内容
中文总结 AI 辅助
提出CasCVS-Net分阶段多任务级联,通过预测边界框和掩码耦合检测、分割与CVS评估,在Endoscapes上训练,超越单任务基线和最先进方法,提升罕见解剖结构定位。
中文摘要 AI 辅助
在腹腔镜胆囊切除术中,对安全关键视图(CVS)的自动化评估既需要识别CVS的三项标准,又需要将解剖结构定位到细小、罕见且常被遮挡的肝囊结构上。基于学习的方法在所用解剖信息上各不相同,从图像级分类到检测、分割或基于图的推理,然而将安全关键的解剖结构进行定位仍是主要瓶颈。我们提出了CasCVS-Net,一种分阶段的多任务级联网络,它联合执行目标检测、语义分割和CVS评估,并在Endoscapes数据集上进行训练。该模型通过预测的解剖结构将任务耦合:预测的边界框引导分割,预测的掩码为CVS分类提供区域级特征,因此在推理时CVS评估仅使用模型预测而非真实标注。为减少这种耦合设置中的优化不稳定性,训练从检测逐步推进到检测-分割,再到完整的三任务级联,随后进行任务级微调。在公开的未见测试集上的评估显示,CasCVS-Net在所有三个任务上均优于匹配的单任务基线,实现了32.0的检测mAP、46.8的语义mIoU、15.3的罕见解剖mIoU和67.2的CVS mAP。它相对于最先进的LG-CVS和SV2LSTG分别提升了6.3%和4.5%的相对CVS mAP,对应4.0和2.9个mAP点。这些结果表明,通过预测的边界框和掩码进行分阶段任务耦合,改善了CVS评估的解剖定位,特别是对于罕见的肝囊结构。
英文摘要
Automated assessment of the Critical View of Safety (CVS) in laparoscopic cholecystectomy requires both recognition of the three CVS criteria and anatomical grounding in small, rare, and often occluded hepatocystic structures. Learning-based methods differ in the anatomical information they use, from image-level classification to detection, segmentation, or graph-based reasoning, yet grounding the safety-critical anatomy remains the main bottleneck. We propose CasCVS-Net, a staged multi-task cascade that jointly performs object detection, semantic segmentation, and CVS assessment, trained on the Endoscapes dataset. The model couples the tasks through predicted anatomy: predicted boxes guide segmentation, and predicted masks provide region-level features for CVS classification, so CVS assessment at inference uses only model predictions rather than ground-truth annotations. To reduce optimisation instability in this coupled setting, training progresses from detection to detection-segmentation and then to the full three-task cascade, followed by task-wise fine-tuning. Evaluation on the public unseen test set shows that CasCVS-Net improves over matched single-task baselines on all three tasks, achieving 32.0 detection mAP, 46.8 semantic mIoU, 15.3 rare-anatomy mIoU, and 67.2 CVS mAP. It outperforms the state-of-the-art LG-CVS and SV2LSTG by 6.3% and 4.5% relative CVS mAP, respectively, corresponding to 4.0 and 2.9 mAP points. These results show that staged task coupling through predicted boxes and masks improves anatomical grounding for CVS assessment, particularly for rare hepatocystic structures.
发表机构
- University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。