发表机构
Texas A&M University(德克萨斯A&M大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MRSeg,一种参数高效的医学图像分割框架,通过联合路由和区域细化实现文本引导的视觉与文本特征自适应,在QaTa-COV19和MosMedData+上取得领先性能。
AI 中文摘要
文本描述可以通过指定要描绘的发现和位置来减少医学图像分割中的歧义。现有的文本引导方法主要改进图像与语言特征交互的位置,但通常在所有图像-文本对之间保留单一的学习更新路径。我们提出MRSeg,一种参数高效的框架,利用每个图像-文本对在密集预测之前引导视觉和文本特征的适应。冻结的ConvNeXt-Tiny和PubMedBERT编码器提供多尺度视觉特征和临床文本标记。联合路由器使用最深的视觉特征和池化文本预测低秩适配器基上的稀疏混合。得到的路由在用于两个视觉尺度和文本的独立适配器库之间共享,协调它们的适应,同时保持特征特定参数分离。区域桥使用文本派生的查询将密集视觉标记聚合为潜在区域,通过自注意力和文本交叉注意力细化这些区域,并将细化后的信息重新分布回特征图。最后,多尺度解码器将细化的语义特征与浅层图像证据结合。在QaTa-COV19和MosMedData+上,MRSeg分别达到90.90/83.32和81.53/68.82的Dice/mIoU,具有7.11M可训练参数和7.60 GFLOPs。代码:此https URL。
英文摘要
Textual descriptions can reduce ambiguity in medical image segmentation by specifying the finding and location to be delineated. Existing text-guided methods mainly improve where image and language features interact but generally retain a single learned update pathway across all image-text pairs. We propose MRSeg, a parameter-efficient framework that uses each image-text pair to route the adaptation of visual and textual features before dense prediction. Frozen ConvNeXt-Tiny and PubMedBERT encoders provide multiscale visual features and clinical text tokens. A joint router uses the deepest visual feature and pooled text to predict a sparse mixture over low-rank adapter bases. The resulting route is shared across separate adapter banks for two visual scales and text, coordinating their adaptation while keeping the feature-specific parameters separate. Region Bridge uses text-derived queries to aggregate dense visual tokens into latent regions, refines these regions through self-attention and text cross-attention, and redistributes the refined information back to the feature maps. Finally, a multiscale decoder combines refined semantic features with shallow image evidence. On QaTa-COV19 and MosMedData+, MRSeg achieves 90.90/83.32 and 81.53/68.82 Dice/mIoU, respectively, with 7.11M trainable parameters and 7.60 GFLOPs. Code: https://github.com/maklachur/MRSeg.
CommentsAccepted at MICCAI 2026 (TIA). Final version to appear in the proceedings