arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向解剖结构无关分割的提示条件通道注意力用于分层特征调制

Prompt-Conditioned Channel Attention for Hierarchical Feature Modulation toward Anatomy-Agnostic Segmentation

Mosharof Hossain, Md Rabiul Islam, Limon Halder, Erchin Serpedin, Md Kamrul Hasan

arXiv 2608.20229首次发表:更新:

发表机构

Khulna University of Engineering & Technology (KUET); Texas A&M University; Imperial College London(库尔纳工程技术大学; 德克萨斯农工大学; 伦敦帝国学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有医学图像分割方法的提示融合局限,提出PCCA机制与PROMISE-Net两种变体,在四个基准测试中较对应基线取得一致IoU提升,为医学图像分割提供可扩展的提示感知分层特征调制框架。

AI 中文摘要

解剖学上合理的分割仍然具有挑战性,原因在于低对比度、模糊的边界以及模态特定的伪影。交互式分割已成为引导特征提取并提升定位能力的有前景策略,尤其适用于结构模糊区域。然而,现有方法通过后期融合整合提示,且缺乏在分层特征表示中实现提示驱动的逐通道调制的明确机制,限制了其捕捉更深层上下文及模态特定变化的能力。为解决这些局限,我们提出提示条件通道注意力(Prompt-Conditioned Channel Attention, PCCA),这是一种新型调制机制,可在编码器-解码器网络中实现语义提示的深度分层整合。PCCA通过池化提取紧凑的通道描述符,将其投影到共享空间,并通过门控激励机制进行融合,以计算感知提示的通道注意力权重。这些权重自适应地重新校准多个网络阶段的特征响应,实现提示条件下、语义丰富的分层表示。在此基础上,我们提出PROMISE-Net,包含两种网络变体:卷积模型(PROMISE-CNN)和基于Transformer的模型(PROMISE-Txformer)。在ISIC-Lesion、Kvasir-Polyp、CAMUS-Cardiac和Kvasir-Instrument基准测试中,将PCCA集成到PROMISE-CNN中,相较于基线U-Net,分别取得了10.4%、8.7%、0.8%和3.4%的相对IoU提升;而PROMISE-Txformer相较于基线UNETR,分别取得了7.6%、23.0%、2.1%和1.1%的对应提升。这些结果表明,在不同架构、成像模态和解剖目标中均实现了一致的改进,确立了PCCA和PROMISE-Net作为可扩展、可泛化的框架,适用于医学图像分割中感知提示的分层特征调制。

英文摘要

Anatomically plausible segmentation remains challenging because of low contrast, ambiguous boundaries, and modality-specific artifacts. Interactive segmentation has emerged as a promising strategy to guide feature extraction and improve localization, particularly in structurally ambiguous regions. However, existing methods integrate prompts through late-stage fusion and lack explicit mechanisms for prompt-driven channel-wise modulation across hierarchical feature representations, limiting their ability to capture deeper contextual and modality-specific variations. To address these limitations, we introduce Prompt-Conditioned Channel Attention (PCCA), a novel modulation mechanism that enables deep, hierarchical integration of semantic prompts within encoder-decoder networks. PCCA extracts compact channel descriptors via pooling, projects them into a shared space, and fuses them through a gated excitation mechanism to compute prompt-aware channel attention weights. These weights adaptively recalibrate feature responses across multiple network stages, enabling prompt-conditioned, semantically enriched hierarchical representations. Building on this, we propose PROMISE-Net, instantiated in two network variants: a convolutional model (PROMISE-CNN) and a transformer-based model (PROMISE-Txformer). Across the ISIC-Lesion, Kvasir-Polyp, CAMUS-Cardiac, and Kvasir-Instrument benchmarks, integrating PCCA into PROMISE-CNN yielded relative IoU gains of 10.4%, 8.7%, 0.8%, and 3.4%, respectively, over the baseline U-Net, while PROMISE-Txformer achieved corresponding gains of 7.6%, 23.0%, 2.1%, and 1.1%, respectively, over the baseline UNETR. These results show consistent improvements across architectures, imaging modalities, and anatomical targets, establishing PCCA and PROMISE-Net as a scalable, generalizable framework for prompt-aware hierarchical feature modulation in medical image segmentation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑