arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAMRI-3D:通过全局体积令牌将SAM2应用于3D MRI分割

SAMRI-3D: Adapting SAM2 for 3D MRI Segmentation with Global Volume Tokens

Zhao Wang, Wei Dai, Hongfu Sun, Craig Engstrom, Shekhar S. Chandra

arXiv 2607.18014首次发表:更新:

发表机构

School of Electrical Engineering and Computer Science, The University of Queensland; School of Engineering, College of Engineering, Science and Environment, University of Newcastle; School of Human Movement and Nutrition Sciences, The University of Queensland(电气工程与计算机科学学院,昆士兰大学; 工程学院,工程、科学与环境学院,纽卡斯尔大学; 人类运动与营养科学学院,昆士兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对MRI低软组织对比度导致边界难观察的问题,提出SAMRI-3D方法,通过冻结图像编码器、微调解码器和内存模块,并引入全局体积令牌,在3D MRI分割上取得高精度,提升了基于SAM的医学模型性能。

AI 中文摘要

诸如Segment Anything Model 2(SAM2)之类的基础模型已经改变了自然图像和视频分割,近期工作开始将它们应用于医学成像。然而,这些改编大多是通用模型,将MRI视为众多模态之一。尽管MRI的低软组织对比度使许多边界在单个切片上难以有效观察到,但大规模、特定于MRI的建模和基准测试仍然有限。我们提出了SAMRI-3D,一种使用SAM2进行3D MRI分割的基准和方法。SAMRI-3D基准是迄今为止最大的仅针对MRI的评估,包含来自34个数据集(27个公开,7个内部)的10392个体积数据,涵盖12个解剖领域和10多种序列,并进行了明确的可见/不可见分割。冻结图像编码器并仅微调轻量级解码器和内存模块,使平均Dice从0.58(零样本SAM2)提高到0.76,显著超过了近期基于SAM的医学模型(SAMed-2为0.69,Medical-SAM2为0.49,SAM-Med3D为0.37)。为了处理不可见边界,我们引入了全局体积令牌(GVT):通过截断符号距离场(TSDF)重建目标训练的持久内存令牌,在推理时被丢弃(零额外成本)。完整的SAMRI-3D模型在所有34个数据集中达到了最佳精度(0.78)和最低方差,并且在8个保留数据集中没有下降(不可见数据为0.79,可见数据为0.78);逐序列分析证实,TSDF目标在切片对比度最弱的地方帮助最大。我们将在本文中发布基准、代码和模型。

英文摘要

Foundation models such as Segment Anything Model 2 (SAM2) have transformed natural-image and video segmentation, and recent work has begun adapting them to medical imaging. These adaptations, however, are largely general-purpose models that treat MRI as one modality among many; large-scale, MRI-specific modelling and benchmarking remain limited, even though MRI's low soft-tissue contrast leaves many boundaries effectively invisible on individual slices. We present SAMRI-3D, a benchmark and method for 3D MRI segmentation with SAM2. The SAMRI-3D benchmark is the largest MRI-only evaluation to date - 10,392 volumes from 34 datasets (27 public, 7 in-house) spanning 12 anatomical domains and 10+ sequences, with explicit seen/unseen splits. Freezing the image encoder and fine-tuning only the lightweight decoder and memory modules raises mean Dice from 0.58 (zero-shot SAM2) to 0.76, surpassing recent SAM-based medical models (SAMed-2 0.69, Medical-SAM2 0.49, SAM-Med3D 0.37) with strong statistical significance. To target invisible boundaries, we introduce Global Volume Tokens (GVT): persistent memory tokens trained with a Truncated Signed Distance Field (TSDF) reconstruction objective that is discarded at inference (zero added cost). This full model, SAMRI-3D, attains the best accuracy (0.78) and lowest variance across all 34 datasets and, uniquely, shows no drop on 8 held-out datasets (0.79 unseen vs. 0.78 seen); per-sequence analysis confirms the TSDF objective helps most where per-slice contrast is weakest. We will release the benchmark, code, and models in this paper.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑