发表机构
College of Information Science and Engineering, Ritsumeikan University(立命馆大学信息科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出SAM+D框架,通过DRLoRA与DSM两个轻量模块,以仅微调约2.8%(SAM)或3.7%(SAM2)参数的方式,实现2D SAM至3D、SAM2至4D的高效分割,在多个基准上取得具竞争力结果。
AI 中文摘要
现有的将SAM等2D基础模型适配至3D体积的方法要么独立处理切片(忽略切片间上下文),要么需要大量架构改动与重新训练。本文提出SAM+D,这是一种参数高效的框架,可将SAM系列模型提升一个空间维度——能基于2D SAM实现3D体积分割,且首次通过参数高效微调实现基于视频的SAM2的端到端4D(3D+时间)时空分割,同时保持绝大多数预训练参数冻结。SAM+D在冻结的Transformer模块中引入两个轻量、模型无关的模块:(1)深度路由LoRA(DRLoRA)专家,带有学习到的路由以实现空间自适应低秩更新;(2)深度移位模块(DSM),用于跨切片特征交换且无额外参数成本。二者共同提供体积级上下文,同时仅微调约2.8%的SAM参数和约3.7%的SAM2参数。我们在两种不同设置下评估SAM+D,每种设置将基础模型提升一个空间维度:3D分割中,SAM(2D→3D)在四个CT基准(KiTS、Pancreas、LiTS、Colon)上评估;4D分割中,SAM2(2D+T→3D+T)在细胞追踪挑战(CTC)数据集(Fluo-N3DH-SIM+)上评估。在两种设置下,SAM+D在单点提示设置下均取得具竞争力或更优的结果,同时使用比现有方法更少的可训练参数,证明SAM+D可泛化至SAM系列架构、目标维度(3D、4D)以及涵盖医学成像和生物场景理解的领域。代码公开可用。
英文摘要
Existing methods for adapting 2D foundation models such as SAM to 3D volumes either process slices independently---ignoring inter-slice context---or require substantial architectural changes and retraining. In this paper, we present \textbf{SAM+D}, a parameter-efficient framework that lifts SAM-family models by one spatial dimension---enabling 3D volumetric segmentation from 2D SAM and, for the first time via parameter-efficient fine-tuning, end-to-end 4D (3D+T) spatiotemporal segmentation from video-based SAM2---while keeping the vast majority of pre-trained parameters frozen. SAM+D introduces two lightweight, model-agnostic modules into frozen transformer blocks: (1)~\textbf{Depth-Routed LoRA (DRLoRA)} experts with learned routing for spatially adaptive low-rank updates, and (2)~\textbf{Depth Shift Modules (DSM)} for cross-slice feature exchange at zero additional parameter cost. Together, they provide volume-level context while tuning only ${\sim}$2.8\% of parameters for SAM and ${\sim}$3.7\% for SAM2. We evaluate SAM+D in two distinct settings, each lifting the base model by one spatial dimension: 3D segmentation, where SAM(2D$\,\to\,$3D) is evaluated on four CT benchmarks (KiTS, Pancreas, LiTS, Colon), and 4D segmentation, where SAM2 (2D+T$\,\to\,$3D+T) is evaluated on a cell tracking challenge (CTC) dataset (Fluo-N3DH-SIM+). In both settings SAM+D achieves competitive or superior results under the single-point prompt setting while using fewer trainable parameters than existing methods, demonstrating that SAM+D generalizes across SAM-family architectures, target dimensionalities (3D, 4D), and domains spanning medical imaging and bio-scene understanding. Code is publicly available at https://github.com/JerrySongCST/SAM-Plus-D.
CommentsAccepted to ECCV2026