arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2501.03565cs.CV

零样本3D医学图像诊断的桥接语义对齐

Bridged Semantic Alignment for Zero-shot 3D Medical Image Diagnosis

  • University of Science and Technology of China(中国科学技术大学)
  • Suzhou Institute for Advanced Research(苏州先进研究所)
  • Stanford University(斯坦福大学)
  • iFlytek Co. Ltd.(iFlytek公司)
  • The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, USTC(中国科学技术大学第一附属医院)

机构由 AI 辅助整理,请以论文原文为准。

Haoran Lai, Zihang Jiang, Qingsong Yao, Rongsheng Wang, Zhiyang He, Xiaodong Tao, Weifu Lv, Wei Wei, S. Kevin Zhou

更新

AI总结:

针对现有视觉-语言对齐方法在零样本3D医学图像诊断中视觉与文本嵌入存在鸿沟的问题,提出桥接语义对齐(BrgSA)框架,通过语义总结和跨模态知识交互缩小差距,在相关数据集上取得最先进性能。

AI中文摘要:

计算机断层扫描等3D医学图像在临床实践中广泛应用,为自动诊断提供巨大潜力。基于监督学习的方法虽取得显著进展,但严重依赖大量人工标注,受限于训练数据可用性及异常类型多样性。视觉-语言对齐(VLA)无需额外标注即可实现零样本学习,是一种有前景的替代方案。然而,我们通过实验发现,现有VLA方法对齐后的视觉和文本嵌入形成两个分离良好的簇,存在需弥合的巨大鸿沟。为此,我们提出桥接语义对齐(BrgSA)框架:首先,利用大语言模型对报告进行语义总结,提取高级语义信息;其次,设计跨模态知识交互模块,以跨模态知识库为语义桥梁,促进两模态间交互,缩小鸿沟并提升对齐效果。为全面评估该方法,我们构建包含15种代表性不足异常的基准数据集,并使用两个现有基准数据集。实验结果表明,BrgSA在公共基准数据集和自定义标注数据集上均取得最先进性能,在代表性不足异常的零样本诊断中实现显著提升。

英文摘要:

3D medical images such as computed tomography are widely used in clinical practice, offering a great potential for automatic diagnosis. Supervised learning-based approaches have achieved significant progress but rely heavily on extensive manual annotations, limited by the availability of training data and the diversity of abnormality types. Vision-language alignment (VLA) offers a promising alternative by enabling zero-shot learning without additional annotations. However, we empirically discover that the visual and textural embeddings after alignment endeavors from existing VLA methods form two well-separated clusters, presenting a wide gap to be bridged. To bridge this gap, we propose a Bridged Semantic Alignment (BrgSA) framework. First, we utilize a large language model to perform semantic summarization of reports, extracting high-level semantic information. Second, we design a Cross-Modal Knowledge Interaction module that leverages a cross-modal knowledge bank as a semantic bridge, facilitating interaction between the two modalities, narrowing the gap, and improving their alignment. To comprehensively evaluate our method, we construct a benchmark dataset that includes 15 underrepresented abnormalities as well as utilize two existing benchmark datasets. Experimental results demonstrate that BrgSA achieves state-of-the-art performances on both public benchmark datasets and our custom-labeled dataset, with significant improvements in zero-shot diagnosis of underrepresented abnormalities.

↑