arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

开放超声基础模型:面向异构场景的稳健分割与临床测量

Open ultrasound foundation model for robust segmentation and clinical measurement across heterogeneous settings

Chao Qin, Fahad Shahbaz Khan, Salman Khan, Sarim Ather, Siddiq Anwar, Rao Muhammad Anwer, Shadab Khan

arXiv 2609.19230首次发表:更新:

发表机构

Mohamed bin Zayed Uni. of Artificial Intelligence; Sheikh Tahnoon Bin Mohammed Medical City (STMC); King’s College Hospital London - Dubai; ADIA Lab(穆罕默德·本·扎耶德人工智能大学; 谢赫·塔赫农·本·穆罕默德医疗城; 伦敦国王学院医院迪拜分院; ADIA实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出开放超声基础模型SonoBase,基于大规模多源数据预训练,在异构场景下分割与临床测量超越现有基线,并具备少样本适应能力。

AI 中文摘要

超声是全世界部署最广泛的影像模态,然而临床AI仍碎片化为狭窄的单任务模型,当设备、操作者或解剖结构变化时,这些模型会失效。在此,我们提出SonoCorpus,一个开放资源,统一了来自53个公共数据集的456,963张图像和1,626,085个专家掩膜,涵盖24个临床应用和17个国家;以及SonoBase,一个在其上预训练的交互式分割基础模型。在引入新器官、设备、操作者和地域的十五个评估数据集上,SonoBase在每个数据集上都优于SAM2、MedSAM2和概念可提示的MedSAM3,并与在同一数据上训练的每数据集专家模型相匹配;在完全外部数据上,其准确率超过了这些基线在各自分布内基准上达到的准确率。由其分割导出的射血分数落在观察者间变异性范围内(误差6.63%),在除颤器候选阈值处的误分类少于任一可提示基线(13%对比18–42%);胎儿头围(1.81毫米)和胎龄(1.2天)误差低于观察者间变异性。在基线完全失败的情况下(每四个测试用例中有一个),SonoBase在81%的此类用例中恢复了可用分割,包括在两个中低收入国家(塞拉利昂和坦桑尼亚)由训练最少的使用者操作的手持探头。五个标注示例可帮助模型适应新环境,且相同的训练协议能很好地迁移到较新的模型(如SAM3),表明优势在于超声特异性预训练而非任何单一架构。为确保可复现性并让社区能够将SonoBase作为平台进行构建,我们发布了所有检查点、优化器状态、数据划分索引、去重哈希和起始代码。

英文摘要

Ultrasound is the most widely deployed imaging modality worldwide, yet clinical AI remains fragmented into narrow single-task models that fail when device, operator, or anatomy changes. Here we present SonoCorpus, an open resource unifying 456,963 images and 1,626,085 expert masks from 53 public datasets spanning 24 clinical applications and 17 countries, and SonoBase, an interactive segmentation foundation model pretrained on it. Across fifteen evaluation datasets introducing new organs, devices, operators, and geographies, SonoBase outperforms SAM2, MedSAM2, and the concept-promptable MedSAM3 on every dataset and matches per-dataset specialist models trained on the same data; on fully external data it exceeds the accuracy these baselines achieve on their own in-distribution benchmarks. Ejection fraction derived from its segmentations falls within inter-observer variability (6.63\% error), with fewer misclassifications at the defibrillator-candidacy threshold than either promptable baseline (13\% versus 18--42\%); fetal head-circumference (1.81~mm) and gestational-age (1.2 days) errors fall below inter-observer variability. Where a baseline fails outright, one in four test cases, SonoBase recovers a usable segmentation in 81\% of them, including on handheld probes operated by minimally trained users in two low- and middle-income countries (Sierra Leone and Tanzania). Five labeled examples can help the model adapt to a new setting, and the identical training protocol transfers well to newer models such as SAM3, locating the advantage in ultrasound-specific pretraining rather than any single architecture. To ensure reproducibility and enable the community to build on SonoBase as a platform, we release all checkpoints, optimizer states, data-split indices, deduplication hashes, and starter code.

CommentsThe PDF includes the Supplementary Information

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑