arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于稀疏手术演示标注的超声机器人气管解剖结构分层学习理解

Learning-based Hierarchical Tracheal Anatomy Understanding from Sparse Surgical Demonstration Annotations for Ultrasound Robots

Hiu Ching Cheung, Wenchao Yue, Zhengran Han, Mingcong Chen, Guanglin Cao, Hongbin Liu, Hongliang Ren

arXiv 2607.22789首次发表:更新:

发表机构

The Chinese University of Hong Kong; Imperial College London; Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences; SLAI Shenzhen Loop Area Institute(香港中文大学; 伦敦帝国学院; 中国科学院香港创新研究院; 深圳环域先进人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对气管切开术定位问题,提出结合YOLOv8n与SAM2的学习框架,经混合训练平衡多方面性能,优于U-Net基线,实现高吞吐量,为机器人辅助气管切开术提供标准化自主手术辅助基础,提升安全性与精度。

AI 中文摘要

气管切开术需要精确确定气管切口位置,传统手动触诊主观且不可靠,超声使用依赖操作者。本文提出基于学习的框架用于分层气管解剖理解,专为超声引导机器人系统设计。采用两阶段感知管道,集成YOLOv8n定位主干和稀疏、提示优化的SAM2解码器,通过混合训练策略确保临床稳健性。实验表明该解耦架构有效平衡泛化、精度和效率,优于U-Net基线,模型吞吐量达6.92 FPS,为标准化自主手术辅助奠定基础,提升机器人辅助气管切开术安全性和精度。

英文摘要

Tracheostomy requires precise localization of the tracheal incision site; however, conventional manual palpation is subjective and often unreliable, while ultrasound utility remains operator-dependent. This work presents a learning-based framework for hierarchical tracheal anatomy understanding, designed specifically for ultrasound-guided robotic systems. We propose a two-stage perception pipeline integrating a YOLOv8n localization backbone with a sparse, prompt-optimized SAM2 decoder to achieve high-fidelity segmentation from sparse surgical annotations. Our hybrid training strategy, bridging curated laboratory data with unconstrained sequences, ensures clinical robustness. Experimental benchmarks demonstrate that this decoupled architecture effectively balances generalization, precision, and efficiency. The YOLOv8n and SAM2 framework achieves a consistent Mean Dice Similarity Coefficient (DSC) of 0.777 across both controlled and generalized domains. This significantly outperforms U-Net baselines, which often suffer from anatomical fragmentation and performance degradation (Generalization DSC $\le$ 0.494). By constraining mask decoding to targeted, sparse regions of interest, our model achieves a throughput of 6.92 FPS, which is vital for closed-loop robotic teleoperation. This study confirms that a robust hierarchical understanding of tracheal anatomy can be derived by coupling lightweight localization with foundation-scale visual models. Our framework establishes a scalable foundation for standardized, autonomous surgical assistance, effectively navigating the variability of real-world ultrasound to enhance the safety and precision of robotic-assisted tracheostomy.

Comments17 pages, 8 figures, accepted at the 2026 International Conference on Cyborg and Bionic Systems (2026ICCBS)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑