发表机构
University of Nottingham Ningbo China; University of Nottingham; University of Lincoln; Jinan University; The Hong Kong Polytechnic University(宁波诺丁汉大学; 诺丁汉大学; 林肯大学; 暨南大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对超声图像分析中任务多样且领域差距大的问题,提出频率引导的多任务路由基础模型FreqDINO++,通过MR-Adapter、F$^2$-Enhancer和TC-Decoder实现参数高效集成与任务协作,在27个临床任务上优于现有方法。
AI 中文摘要
超声图像分析在癌症筛查和产前诊断中起着至关重要的作用,然而全面评估需要联合处理病灶分割和良恶性分类等任务。尽管最近的视觉基础模型已展现出显著的通用表示能力,但将其潜力应用于超声领域受到与自然图像之间巨大领域差距的制约。现有方法通常针对孤立任务微调重型视觉编码器,导致大量计算开销,同时忽视了异构任务之间的潜在共性。在本工作中,我们提出了FreqDINO++,一种用于通用超声分析的频率引导多任务路由视觉基础模型。我们首先引入多任务路由适配器(MR-Adapter)以支持任务通用和任务特定知识的参数高效集成,然后设计了频率感知特征增强器(F$^2$-Enhancer)来捕获超声图像丰富的多尺度频率特征,并设计了任务对齐协作解码器(TC-Decoder)通过全局-局部令牌交互促进密集预测任务和全局预测任务之间的协作。在大规模多任务和外部单任务超声基准上的大量实验表明,FreqDINO++在27个不同的临床任务场景中持续优于强基线和最近的基础模型,同时展现出对未见数据的有前景的泛化能力。代码见该https URL。
英文摘要
Ultrasound image analysis plays a crucial role in cancer screening and prenatal diagnosis, yet comprehensive assessment requires jointly addressing tasks such as lesion segmentation and benign-malignant classification. While recent vision foundation models have shown remarkable universal representations, unlocking their potential for ultrasound is bottlenecked by the considerable domain gap from natural images. Existing methods typically fine-tune heavy vision encoders for isolated tasks, incurring substantial computational overhead while overlooking the underlying commonalities across heterogeneous tasks. In this work, we propose FreqDINO++, a frequency-guided multi-task routing vision foundation model for universal ultrasound analysis. We first introduce a Multi-task Routing Adapter (MR-Adapter) to support parameter-efficient integration of task-common and task-specific knowledge, a Frequency-aware Feature Enhancer (F$^2$-Enhancer) is then designed to capture the rich multi-scale frequency characteristics of ultrasound images, and a Task-aligned Collaborative Decoder (TC-Decoder) is devised to promote collaboration between dense and global prediction tasks through global-local token interaction. Extensive experiments on large-scale multi-task and external single-task ultrasound benchmarks demonstrate that FreqDINO++ consistently outperforms strong baselines and recent foundation models across 27 diverse clinical task scenarios, while also showing promising generalization to unseen data. The code is at https://github.com/MingLang-FD/FreqDINO-Plus.
CommentsAccepted by TBME