arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18731cs.CVcs.AI

少数案例足矣:对MedSAM3的标注高效LoRA微调的实证研究

A Few Cases Are All You Need: An Empirical Study of Annotation-Efficient LoRA Fine-Tuning of MedSAM3

  • Norwegian University of Science and Technology(挪威科技大学)
  • St. Olavs Hospital, Trondheim University Hospital(圣奥拉夫斯医院(特隆赫姆大学医院))

机构由 AI 辅助整理,请以论文原文为准。

Sachin Dudda Nagaraju, Bendik Skarre Abrahamsen, Ashkan Moradi, Mattijs Elschot

AI总结:

该研究针对MedSAM3,用LoRA微调实现仅10个标注案例即可让医学图像分割达到临床可用性能,在腹部器官和心脏分割任务上优于部分专业工具,且训练速度更快。

AI中文摘要:

医学图像分割对于治疗规划、疾病评估等临床流程至关重要。虽然TotalSegmentator和MRSegmentator等专业工具性能强劲,但它们的训练需要大量标注数据集。医学基础模型通过大规模预训练减轻了新任务的标注负担,是一种颇具前景的替代方案,但其零样本性能仍有限。通过低秩适配(LoRA)实现的参数高效适配,仅需少量可训练参数即可完成高效专业化,但一个关键问题依然存在:需要多少专家标注的案例才能达到临床可用的分割性能?我们针对这一问题,将MedSAM3通过LoRA适配,在CT和MRI影像中针对肝脏、肾脏、脾脏、胆囊和胰腺这五个腹部器官,仅使用1、2、5和10个标注案例进行训练,并在AMOS22数据集上评估性能。仅用10个案例时,模型的性能可与用数量级更多数据训练的专业系统相媲美。值得注意的是,这其中包括可靠的胆囊分割(CT影像的Dice系数为0.68,MRI影像的Dice系数为0.59),而现有工具在该任务上几乎完全失效(Dice系数为0.0004);同时,对于肝脏、肾脏和脾脏,其性能与MRSegmentator的差距保持在5%至10%之间,而使用的标注量减少了100倍以上。此外,在全心脏分割数据集上的外部验证表明,该方法可扩展至心脏分割这一超出TotalSegmentator(MRI)和MRSegmentator适用范围的用例,仅用10个标注案例即可达到具有竞争力的左心室(LV)分割性能。每个器官的训练仅需在单个GPU上花费3至5小时,比nnU-Net快约2至3倍。这些发现表明,10个标注案例足以实现临床可用的分割效果,有效减少了图像标注和训练时间方面的瓶颈。

英文摘要:

Medical image segmentation is essential for clinical workflows such as treatment planning and disease assessment. While specialist tools like TotalSegmentator and MRSegmentator achieve strong performance, they require large annotated datasets for training. Medical foundation models offer a promising alternative through large-scale pretraining that reduces the annotation burden for new tasks, but zero-shot performance remains limited. Parameter-efficient adaptation via Low-Rank Adaptation (LoRA) enables efficient specialization with few trainable parameters, but a key question remains: how many expert-annotated cases are needed to achieve clinically useful segmentation performance? We address this by adapting MedSAM3 with LoRA for five abdominal organs (liver, kidneys, spleen, gallbladder, and pancreas) in CT and MRI using only 1, 2, 5, and 10 annotated cases, evaluating on AMOS22 dataset. With just 10 cases, models achieve performance competitive with specialist systems trained on orders of magnitude more data. Notably, this includes reliable gallbladder segmentation (Dice 0.68 CT, 0.59 MRI) where existing tools fail almost completely (Dice 0.0004), while remaining within 5--10% of MRSegmentator for liver, kidneys, and spleen using over 100 times fewer annotations. Furthermore, external validation on the Whole Heart Segmentation dataset shows that the approach extends to cardiac segmentation, a use case beyond the scope of TotalSegmentator (MRI) and MRSegmentator, achieving competitive left ventricle (LV) performance with only 10 annotated cases. Training requires only3--5,hours per organ on a single GPU, approximately 2--3 times faster than nnU-Net. These findings suggest that ten annotated cases are sufficient for clinically useful segmentation, effectively reducing bottlenecks for both image annotation and training time.

↑