发表机构
Harbin Institute of Technology (Shenzhen); NingBo No.2 Hospital(哈尔滨工业大学(深圳); 宁波第二医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对文本引导医学分割中模型组件紧密耦合问题,提出可转移骨干层次适配器框架BTHA,通过稳定特征级接口、分层监督策略和自适应门控语义引导适配器,有效提升分割效果且计算开销适度。
AI 中文摘要
文本引导的医学图像分割利用临床语义来改善病变描绘,但现有许多模型将跨模态融合、监督和解码器设计绑定到特定任务架构中。这种紧密耦合使得语言引导模块难以在异构视觉和文本骨干网络中重用,且编码器对改变时通常需重新设计网络。本文提出了BTHA,一种用于文本引导医学图像分割的可转移骨干层次适配器框架。它围绕稳定的特征级接口构建,通过保持形状的适配器注入语义引导,同时维持解码器端张量收缩。为使该接口有效,引入分层粗细监督策略。还设计了尺度自适应门控语义引导适配器。跨多种骨干网络的评估表明相同适配器和监督设计有效,在四个公共数据集上的实验进一步证明BTHA以适度计算开销改进了强大的文本引导基线。
英文摘要
Text-guided medical image segmentation leverages clinical semantics to improve lesion delineation, yet many existing models bind cross-modal fusion, supervision, and decoder design into a task-specific architecture. Such tight coupling makes it difficult to reuse language guidance modules across heterogeneous vision and text backbones, and often requires redesigning the network when the encoder pair changes. This paper presents BTHA, a backbone-transferable hierarchical adapter framework for text-guided medical image segmentation. BTHA is built around a stable feature-level interface: given multi-scale visual features and a text representation, it injects semantic guidance through shape-preserving adapters while maintaining the decoder-side tensor contract. To make this interface effective, we introduce a Hierarchical Coarse-to-Fine Supervision Strategy that decomposes learning into global image-text alignment, multi-scale auxiliary localization, and boundary-aware final mask refinement. We further design a Scale-Adaptive Gated Semantic Guidance (SAGSG) adapter, where resolution-specific gates adaptively control textual injection and channel recalibration suppresses redundant cross-modal responses. Evaluations across diverse vision and text backbones show that the same adapter and supervision design remains effective across convolutional and transformer-based visual encoders as well as different language encoders. Experiments on four public datasets further demonstrate that BTHA improves strong text-guided baselines with modest computational overhead.