arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ThyCLIPNet:一种BiomedCLIP引导的轻量级注意力增强DeepLabV3+框架,用于稳健的甲状腺结节分割

ThyCLIPNet: A BiomedCLIP-Guided Lightweight Attention-Enhanced DeepLabV3+ Framework for Robust Thyroid Nodule Segmentation

Tasnim Jahan, Md Easin Arafat, Swakkhar Shatabda

arXiv 2610.04743首次发表:更新:

发表机构

United International University; Eötvös Loránd University; BRAC University(联合国际大学; 厄特沃什·罗兰大学; BRAC大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ThyCLIPNet提出一种轻量级语义引导的编码器-解码器框架,结合BiomedCLIP视觉引导与多尺度CNN,在多个甲状腺超声数据集上实现稳健高效的分割。

AI 中文摘要

准确的甲状腺超声分割常常受到低对比度、斑点噪声和边界不清晰的挑战。尽管近期方法提高了分割精度,但许多方法依赖资源密集型架构,或缺乏多尺度特征与全局生物医学视觉引导的显式集成。在本文中,我们提出了ThyCLIPNet,一种轻量级语义引导的混合编码器-解码器框架,将BiomedCLIP衍生的生物医学语义引导集成到轻量级多尺度CNN分割流程中。编码器集成了带有高效通道注意力的MobileNetV2,而空洞空间金字塔池化和自定义卷积块注意力模块丰富了瓶颈特征。解码器结合了层次化跳跃连接和轻量级注意力细化,以及BiomedCLIP引导的门控融合路径,该路径将仅视觉的全局生物医学嵌入投影到解码器特征空间,并通过语义-局部融合和空间门控选择性地集成它们。据我们所知,ThyCLIPNet是首批使用BiomedCLIP的视觉编码器单独进行仅图像全局语义引导(无需文本提示)的轻量级甲状腺超声分割框架之一。在TG3K、TN3K、DDTI和PKTN上的实验分别实现了96.22%、87.58%、84.73%和80.70%的Dice相似系数;92.72%、77.91%、73.51%和67.64%的交并比分数;以及3.75、16.38、18.23和10.86的95百分位Hausdorff距离。ThyCLIPNet使用8.55M参数和22.99G FLOPs。总体而言,结果支持将全局生物医学语义引导与轻量级多尺度CNN表示相结合,以实现稳健且计算高效的甲状腺超声分割。源代码:此https URL。[摘要因arXiv要求缩短,完整摘要见PDF。]

英文摘要

Accurate thyroid ultrasound segmentation is often challenged by low contrast, speckle noise, and unclear boundaries. Although recent methods have improved segmentation accuracy, many rely on resource-intensive architectures or lack explicit integration of multiscale features with global biomedical visual guidance. In this paper, we introduce ThyCLIPNet, a lightweight semantic-guided hybrid encoder-decoder framework that integrates BiomedCLIP-derived biomedical semantic guidance into a lightweight multi-scale CNN segmentation pipeline. The encoder integrates MobileNetV2 with efficient channel attention, while atrous spatial pyramid pooling and a custom convolutional block attention module enrich bottleneck features. The decoder combines hierarchical skip connections and lightweight attention refinement with a BiomedCLIP-guided gated fusion pathway that projects vision-only global biomedical embeddings into decoder feature space and selectively integrates them through semantic-local fusion and spatial gating. To the best of our knowledge, ThyCLIPNet is among the first lightweight thyroid ultrasound segmentation frameworks to use BiomedCLIP's vision encoder alone for image-only global semantic guidance without text prompting. Experiments on TG3K, TN3K, DDTI, and PKTN achieve dice similarity coefficients of 96.22%, 87.58%, 84.73%, and 80.70%; intersection over union scores of 92.72%, 77.91%, 73.51%, and 67.64%; and 95th-percentile hausdorff distances of 3.75, 16.38, 18.23, and 10.86, respectively. ThyCLIPNet uses 8.55M parameters and 22.99G FLOPs. Overall, the results support integrating global biomedical semantic guidance with lightweight multi-scale CNN representations for robust and computationally efficient thyroid ultrasound segmentation. Source code: https://github.com/Tasnim-Jahan/ThyCLIPNet. [Abstract shortened for arXiv. See PDF for full abstract.]

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑