arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11335cs.CVcs.AIeess.IV

用于临床文本引导的医学图像分割的双域跨模态解码

Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation

  • Texas A&M University(德克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

Md Maklachur Rahman, Tracy Hammond

AI总结:

本文针对临床文本引导医学图像分割忽略频率内容的问题,提出双域跨模态解码(DD-CMD)方法,在两个公开数据集上均优于现有最强基线。

AI中文摘要:

临床文本可缩小需分割的目标范围,但近期文本引导的设计方案侧重空间对齐,却忽略了决定纹理与边界的频率内容。本文针对临床文本引导的肺部感染分割,提出双域跨模态解码(DD-CMD),在解码阶段整合两种互补形式的语言引导。在空间域,文本引导空间交叉注意力(TGSA)将多尺度视觉标记与文本语义对齐,并通过门控残差融合更新特征;在频率域,频谱-文本自适应调制(STAM)应用二维离散余弦变换(2D DCT)计算可学习的频带能量统计量,预测文本条件下的FiLM参数,以重新校准解码器通道,实现感知频率的解码。DD-CMD将TGSA与STAM嵌入从7×7到56×56的粗到细解码器,通过轻量两阶段细化模块恢复全分辨率掩码。在QaTa-COV19与MosMedData+数据集上的实验显示,DD-CMD分别取得91.46%的戴斯系数(Dice)、84.26%的均值交并比(mIoU)与81.95%的Dice、69.42%的mIoU,相比最强的现有基线方法,平均提升了1.96的Dice与2.67的mIoU。代码:this https URL。

英文摘要:

Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries. We propose Dual-Domain Cross-Modal Decoding (DD-CMD) for clinical text-guided pulmonary infection segmentation, integrating two complementary forms of language guidance during decoding. In the spatial domain, Text-Guided Spatial Cross-Attention (TGSA) aligns multi-scale visual tokens with text semantics and updates features through gated residual fusion. In the frequency domain, Spectral-Text Adaptive Modulation (STAM) applies a 2D DCT to compute learnable band-energy statistics and predicts text-conditioned FiLM parameters to recalibrate decoder channels for frequency-aware decoding. DD-CMD embeds TGSA and STAM into a coarse-to-fine decoder (7x7 to 56x56) and restores full-resolution masks using a lightweight two-stage refinement module. Experiments on QaTa-COV19 and MosMedData+ show that DD-CMD achieves 91.46% Dice / 84.26% mIoU and 81.95% Dice / 69.42% mIoU, respectively, with average gains of +1.96 Dice and +2.67 mIoU over the strongest prior baselines. Code: https://github.com/maklachur/DD-CMD.

补充信息

↑