发表机构
National Taiwan University of Science and Technology; National Yang Ming Chiao Tung University(国立台湾科技大学; 国立阳明交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对超声中针尖定位的挑战,提出STUNet-Fusion时空框架,通过三通道运动融合与ResNet-34、ConvLSTM和U-Net实现亚像素精度,显著提升鲁棒性。
AI 中文摘要
超声中的针尖定位仍然具有挑战性,因为针可能显得微弱、不连续或部分不可见,而成像伪影和解剖结构可能产生相似的响应。为了解决这个问题,我们提出了STUNet-Fusion,一种用于超声视频中针尖定位的时空框架。所提出的方法将输入构建为三通道时空融合张量,包括灰度外观、基于网格的运动特征和原始帧差。共享的ResNet-34编码器提取空间特征,ConvLSTM整合时间依赖性,U-Net解码器重建密集概率热图。最终坐标通过soft-argmax操作提取,以实现亚像素定位精度。实验结果表明,与传统基线相比,这种时空融合策略显著提高了定位鲁棒性。
英文摘要
Needle-tip localization in ultrasound remains challenging because the needle may appear weak, discontinuous, or partially invisible, while imaging artifacts and anatomical structures can produce similar responses. To address this problem, we propose STUNet-Fusion, a spatiotemporal framework for needle-tip localization in ultrasound videos. The proposed method formulates the input as a tri-channel spatio-temporal fusion tensor, comprising grayscale appearance, grid-based motion feature, and raw frame difference. A shared ResNet-34 encoder extracts spatial features, ConvLSTM integrates temporal dependencies, and a U-Net decoder reconstructs a dense probability heatmap. The final coordinates are extracted via a soft-argmax operation to achieve sub-pixel localization accuracy. Experimental results demonstrate that this spatiotemporal fusion strategy significantly improves localization robustness compared to conventional baselines.