发表机构
The University of Sydney; Shanghai Jiao Tong University; Northeastern University(悉尼大学; 上海交通大学; 东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PureVision框架,通过几何监督学习(PureEyes)和解剖引导证据聚合(PureNeurons)改善医学VLM对多表型病变的解读,在三个数据集上验证了有效性。
AI 中文摘要
医学视觉语言模型(VLMs)在临床图像解读方面展现出越来越大的潜力。然而,这些模型在处理多表型病变时仍然存在困难,这类病变的诊断需要联合评估多种病理表型。现有的视觉-语言对齐方法产生的视觉表示无法保留解剖层次结构和表型亚类之间的关系。这源于它们依赖语义监督,缺乏几何约束来在视觉嵌入空间中保留这些关系。此外,病变相关的解剖和表型表示的稀疏性使得医学VLM难以捕获重要的诊断证据。为解决这些局限性,我们提出了PureVision,一种用于医学VLM中多表型病变解读的几何监督视觉表示学习框架。它结合了几何监督表示学习模块PureEyes和解剖引导的证据聚合模块PureNeurons。PureEyes通过编码解剖层次结构和表型亚类关系的理想空间分布提供几何监督。PureNeurons将视觉表示投影到学习到的潜在空间中,利用其位置选择性地聚合病变特定的解剖和表型证据。在LIDC-IDRI、CBIS-DDSM和3DReasonKnee上的实验表明,PureVision在视觉问答和放射学报告生成中改善了病变定位和表型表征。代码可在以下网址获取:this https URL。
英文摘要
Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation. However, these models still struggle to interpret multi-phenotype lesions whose diagnosis requires the joint assessment of multiple pathological phenotypes. Existing vision-language alignment methods produce visual representations that fail to preserve anatomical hierarchies and relationships among phenotypic subclasses. This stems from their reliance on semantic supervision, which lacks geometric constraints to preserve these relationships in the visual embedding space. Moreover, the sparsity of lesion-related anatomical and phenotypic representations makes it difficult for medical VLMs to capture important diagnostic evidence. To address these limitations, we propose \textbf{PureVision}, a geometry-supervised visual representation learning framework for multi-phenotype lesion interpretation in medical VLMs. It combines a geometry-supervised representation learning module, \textbf{PureEyes}, and an anatomy-guided evidence aggregation module, \textbf{PureNeurons}. PureEyes provides geometric supervision through ideal spatial distributions that encode anatomical hierarchies and phenotypic subclass relationships. PureNeurons projects visual representations into the learned latent space, using their positions to selectively aggregate lesion-specific anatomical and phenotypic evidence. Experiments on \textit{LIDC-IDRI}, \textit{CBIS-DDSM}, and \textit{3DReasonKnee} demonstrate that PureVision improves lesion grounding and phenotype characterization in visual question answering and radiology report generation. Code is available at: https://anonymous.4open.science/r/purevision-06C2.