arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越分类:结构化监督使视觉证据与医学语义对齐

Beyond Classification: Structured Supervision Aligns Visual Evidence with Medical Semantics

Hexiang Bai, Hanyang Xu, Xiaoxue Li, Xiaoliang Wu, Shangde Gao, Hongxia Xu, Ke Liu

arXiv 2609.05937首次发表:更新:

发表机构

Zhejiang University; The Second Affiliated Hospital Zhejiang University; Transvascular Implantation Devices Research Institute and State Key Laboratory of Transvascular Implantation Devices; Henan Normal University; University of Southampton; School of Software Technology Zhejiang University(浙江大学; 浙江大学医学院附属第二医院; 经血管植入器械研究院及经血管植入器械全国重点实验室; 河南师范大学; 南安普顿大学; 浙江大学软件学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对医学图像ViT预训练中的空间坍缩问题,本文系统比较三种结构化监督,发现图像-文本跨模态对齐最优,能增强特征表示并提升下游分类性能与可解释性。

AI 中文摘要

视觉Transformer(ViTs)在医学图像分析中展现出巨大潜力。然而,通过全局图像分类进行的标准预训练存在空间坍缩问题,即模型严重依赖背景捷径而非定位关键的前景病灶。为克服这一局限并使视觉证据与精确的医学语义对齐,我们系统地研究了替代性预训练方法。具体而言,我们评估了三种独立形式的结构化监督:通过图自监督实现的拓扑先验、通过分割实现的密集像素级约束,以及通过图像-文本对实现的跨模态语义 grounding。值得注意的是,我们的实证分析揭示,虽然这三种结构化监督形式均能成功缓解全局池化瓶颈并将视觉注意力引导至前景区域,但图像-文本对齐取得了最优性能。通过嵌入高维诊断逻辑,跨模态方法不仅将注意力锚定在精确的视觉证据上,还实现了深刻的抽象推理。大量实验表明,这种语义丰富的预训练从根本上增强了模型的特征表示。因此,当针对下游临床分类任务进行微调时,我们的模型实现了卓越的准确性,并生成了聚焦于真实病理特征的高度可解释的注意力图,大幅优于普通分类基线。

英文摘要

Vision Transformers (ViTs) have shown immense potential in medical image analysis. However, standard pre-training via global image classification suffers from spatial collapse, where models rely heavily on background shortcuts rather than localising critical foreground lesions. To overcome this limitation and align visual evidence with precise medical semantics, we systematically investigate alternative pre-training paradigms.Specifically, we evaluate three independent forms of structured supervision: topological priors via graph self-supervision, dense pixel-level constraints via segmentation, and cross-modal semantic grounding via image-text pairs. Notably, our empirical analysis reveals that while all three forms of structured supervision successfully alleviate the global pooling bottleneck and steer visual attention towards foreground regions, image-text alignment achieves the most superior performance. By embedding high-dimensional diagnostic logic, the cross-modal approach not only anchors attention on precise visual evidence but also enables profound abstract reasoning. Extensive experiments demonstrate that this semantically enriched pre-training fundamentally enhances the model's feature representation. Consequently, when fine-tuned for downstream clinical classification tasks, our models achieve superior accuracy and yield highly interpretable attention maps focused on true pathological features, vastly outperforming vanilla classification baselines.

Comments22 pages,7 figures,Accepted to BMVC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑