arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SegBanana:将统一多模态模型引导为医学分割器

SegBanana: Steering Unified Multimodal Models into Medical Segmenters

Xiaoye Liang, Ye Yan, Mingze Yin, Shikun Feng, Mai Xu, Haiguang Liu, Lai Jiang, Yiheng Zhu

arXiv 2609.34235首次发表:更新:

发表机构

Beihang University; Zhongguancun Academy; Zhejiang University(北京航空航天大学; 中关村学院; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SegBanana提出首个免训练的智能体视觉生成框架,利用冻结统一多模态模型配合解剖知识检索与质量批判,在八个医学分割数据集上平均mDice达77.45%,显著超越现有基线。

AI 中文摘要

医学图像分割在实际部署中仍然具有挑战性,因为模型往往难以泛化到训练数据所覆盖的分布之外,且高质量的像素级标注通常无法用于适配。受大语言模型跨任务迁移能力的启发,我们研究了统一多模态模型(UMMs)能否将其预训练的视觉理解、推理和生成能力迁移到医学图像分割,而无需针对特定任务的后训练。通过将分割重新定义为结构化视觉生成,我们发现前沿UMMs(例如Nano Banana)在多种临床场景中已展现出基础的分割能力,但在需要专门解剖学或领域特定知识的挑战性任务上仍存在困难。我们进一步表明,这些局限性可以通过结合上下文示例中的视觉解剖知识、通过重复采样扩展候选解决方案以及通过有针对性的自我修正来优化次优预测,从而得到有效缓解。基于这些观察,我们提出了SegBanana,据我们所知,这是首个用于免训练医学图像分割的智能体视觉生成框架。SegBanana以冻结的UMM作为核心生成模型,并辅以解剖感知知识检索和比较质量批判,以释放其潜在的分割能力。一个状态感知多模态控制器维护结构化状态并迭代编排这些工具,反复优化中间预测以生成更高质量的掩膜。在八个医学分割数据集上,SegBanana达到了77.45%的平均mDice,比代表性的通用模型(SAM3和SegGPT)和医学专用模型(BiomedParse和MedSAM3)基线高出至少14.93个百分点,同时对域外视觉支持保持鲁棒性。

英文摘要

Medical image segmentation remains challenging in practical deployment, as models often struggle to generalize beyond the distributions covered by their training data and high-quality pixel-level annotations are typically unavailable for adaptation. Inspired by the cross-task transferability of large language models, we investigate whether unified multimodal models (UMMs) can transfer their pretrained visual understanding, reasoning, and generation capabilities to medical image segmentation without task-specific post-training. By recasting segmentation as structured visual generation, we find that frontier UMMs (e.g., Nano Banana) already exhibit basic segmentation capabilities across diverse clinical scenarios, but still struggle with challenging tasks requiring specialized anatomical or domain-specific knowledge. We further show that these limitations can be effectively mitigated by incorporating visual anatomical knowledge from in-context exemplars, expanding candidate solutions through repeated sampling, and refining suboptimal predictions via targeted editing.Motivated by these observations, we propose SegBanana, to our knowledge, the first agentic visual generation framework for training-free medical image segmentation. SegBanana builds on a frozen UMM as the core generative model, augmented with Anatomy-Aware Knowledge Retrieval and Comparative Quality Critique to unlock its potential segmentation capability. A State-Aware Multimodal Controller maintains structured state and iteratively orchestrates these tools, repeatedly refining intermediate predictions toward higher-quality masks. Across eight medical segmentation datasets, SegBanana achieves an average mDice of 77.45%, outperforming representative generalist (SAM3 and SegGPT) and medical-specific (BiomedParse and MedSAM3) baselines by at least 14.93 points, while remaining robust to out-of-domain visual supports.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑