arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于弱监督的主动学习在手术视频高效标注中的应用

Active Learning for Efficient Annotation of Surgical Videos with Weak Supervision

Manasa Dendukuri, Matjaz Jogan, Daniel A. Hashimoto, Guiqiu Liao

arXiv 2607.13237首次发表:更新:

发表机构

University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对腹腔镜视频标注难题,提出人在回路知识获取框架,结合主动学习与双损失优化,利用基础模型及互补训练目标生成CAMs,迭代提出伪掩码引导标注,减少标注工作量50%,实现模型开发可扩展性及高效知识获取。

AI 中文摘要

腹腔镜视频的精确时空标注既耗时又需要专业知识。我们提出了一种人在回路的知识获取框架,将主动学习与双损失优化相结合,显著减少手术领域中物体自动定位和分割所需的标注工作量。该方法利用基础模型,通过两个互补的训练目标从视频中生成时间一致的类激活映射(CAMs):针对弱标注数据的视频级工具存在标签的弱监督损失,以及通过主动学习获得的人工校正标注上的图像级掩码损失。我们的流程迭代地提出伪掩码,引导专家注释器完善模型先前捕获的知识,而不是一开始就需要密集的像素级标注。实验表明,与完全手动标注相比,我们的框架在训练结束时将手术视频标注工作量减少了50%。该框架无需一开始就有大量完全标注的数据集,实现了手术工具分割模型开发的可扩展性。这种迭代的人在回路细化支持以最少的专家输入进行高效的知识获取,为将工具分割扩展到更大、更多样化的数据集和现实世界临床环境提供了实用且可部署的策略。

英文摘要

Precise spatial-temporal annotation of laparoscopic videos is time-consuming and requires expert knowledge. We propose a human-in-the-loop knowledge acquisition framework that combines active learning with dual-loss optimization to significantly reduce the annotation effort needed for automatic localization and segmentation of objects in the surgical field. Our method employs a foundation model to generate temporally consistent class activation maps (CAMs) from video using two complementary training objectives: a weak supervision loss on video-level tool presence labels for weakly annotated data, and an image-level mask loss on human-corrected annotations obtained through active learning. Rather than requiring dense pixel-level annotation upfront, our pipeline iteratively proposes pseudo-masks that guide the expert annotator to refine the knowledge previously captured by the model. We demonstrate that our framework reduces the effort of surgical video annotation by 50% by the end of training in comparison to fully manual annotation. Through eliminating the need for large, fully annotated datasets from the start, this framework enables scalability to the development of surgical tool segmentation models. This iterative human-in-the-loop refinement supports efficient knowledge acquisition with minimal expert input, providing a practical and deployable strategy for expanding tool segmentation to larger, more diverse datasets and real-world clinical settings.

CommentsAccepted to IPCAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑