arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超级通才:通过通才-专才协同实现全面准确的医学图像理解

Super-Generalist: Towards Comprehensive and Accurate Medical Image Understanding via Generalist-Specialist Synergy

Shaoteng Zhang, Weiwei Cao, Wanxing Chang, Yutong Xie, Kai Cao, Zaiyi Liu, Yu Shi, Tingbo Liang, Qi Zhang, Ling Zhang, Yong Xia, Jianpeng Zhang

arXiv 2607.09135首次发表:更新:

发表机构

DAMO Academy, Alibaba Group; Hupan Lab; Department of Radiology, Ningbo No. 2 Hospital; Department of Radiology, Guangdong Provincial People’s Hospital; Department of Radiology, Shanghai Institution of Pancreatic Disease; Department of Radiology, Shengjing Hospital of China Medical University; The First Affiliated Hospital, Zhejiang University School of Medicine; Mohamed bin Zayed University of Artificial Intelligence(达摩院,阿里巴巴集团; 湖畔实验室; 宁波市第二医院放射科; 广东省人民医院放射科; 上海胰腺疾病研究所放射科; 中国医科大学附属盛京医院放射科; 浙江大学医学院附属第一医院; 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在通过通才-专才协同实现全面准确的医学图像理解。提出SuG框架,整合多分割专家空间先验进行视觉-语言对齐,利用病变掩码校准注意力。在多CT基准测试中评估,其在疾病诊断任务中性能领先,超越专才模型,有强大病变定位能力。

AI 中文摘要

医学图像需要全面准确的解读以辅助各种临床病症的诊断。近期的视觉-语言通才模型任务覆盖广且有零样本能力,但缺乏精细解剖和病变认知。监督式专才模型在特定任务上表现出色,但缺乏跨疾病和解剖结构的泛化能力。本文提出SuG框架,将通才视觉-语言学习与专才目标统一,通过整合多分割专家的空间先验进行专才增强的视觉-语言对齐,利用病变掩码校准视觉注意力。在多个胸部和腹部CT基准测试上评估,SuG在多种疾病诊断任务中取得了领先性能,在关键肿瘤诊断基准测试中超越专才模型,还展示了强大的病变定位能力。

英文摘要

Medical images require comprehensive and accurate interpretation to support the diagnosis of diverse clincial conditions. Recent vision-language generalist models offer broad task coverage and promising zero-shot capabilities, yet often lack fine-grained anatomical and lesion awareness for reliable diagnosis and spatial interpretability. In contrast, supervised specialist models achieve strong performance on specific tasks but typically lack generalization across diseases and anatomies. In this work, we present SuG, a Super-Generalist framework that unifies generalist vision-language learning with specialist objectives, enabling both broad generalization and specialist-level diagnostic capability. We perform specialist-enhanced vision-language alignment in SuG by incorporating spatial priors from multiple segmentation experts, including anatomy, class-specific lesion and class-agnostic lesion segmentors that captures lesions beyond anatomies annotated during training. To improve lesion grounding capability, we leverage lesion masks as spatial priors to calibrate text-conditioned visual attention, encouraging disease-related semantics to focus on clinically relevant regions. We evaluate SuG on extensive chest and abdominal CT benchmarks, including CT-RATE, Merlin, MedVL-CT69K, and several in-house tumor datasets. SuG achieves state-of-the-art performance across a wide range of disease diagnosis tasks and surpasses specialist models on several critical tumor diagnosis benchmarks. Furthermore, SuG demonstrates strong lesion grounding capability, including robust generalization to lesion types lacking class-specific supervision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑