发表机构
The University of Sydney; Shanghai Jiao Tong University; Northeastern University; Jilin University(悉尼大学; 上海交通大学; 东北大学; 吉林大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出$Δ$Representation框架,通过几何监督和反事实推理学习医学VLM的病理表型增量表示,提升病变定位与表型表征准确性。
AI 中文摘要
医学视觉语言模型(VLMs)在放射学图像解读方面展现出日益增长的潜力。医学VLMs将放射学图像编码为视觉表示,以捕获用于诊断的解剖和表型信息。现有方法通过语义引导的表示对齐来改进病理表型表示。然而,病理表型表现为叠加在正常解剖结构之上的病变特异性视觉变化。这种语义对齐方法无法建模相对于相应正常解剖表示的表型特异性增量。为解决这一差距,我们提出了\ extbf{$Δ$Representation},一种基于反事实推理的医学VLM视觉表型表示学习框架。它包含\ extbf{BaseAnatomy},一个几何监督的表示学习模块,以及\ extbf{$Δ$Phenotype},一个反事实增量表示学习模块。BaseAnatomy通过跨解剖结构和内部的空间关系提供细粒度的几何监督。$Δ$Phenotype计算病变表示与其对应正常解剖表示之间的表示增量,并监督与相同表型相关的增量在表示空间中聚类。在\ extit{ReXGroundingCT}和\ extit{LIDC-IDRI}上的实验表明,$Δ$Representation有效结构化病理表型表示,并提高医学VLMs中病变定位和表型表征的准确性。代码可在该https URL获取。
英文摘要
Medical vision-language models (VLMs) have shown increasing potential for radiological image interpretation. Medical VLMs encode radiological images into visual representations that capture both anatomical and phenotypic information for diagnosis. Existing approaches improve pathological phenotype representations through semantic-guided representation alignment. However, pathological phenotypes arise as lesion-specific visual changes superimposed on underlying normal anatomy. Such semantic alignment approaches fail to model the phenotype-specific increment relative to the corresponding normal anatomical representation. To address this gap, we propose \textbf{$Δ$Representation}, a visual phenotype representation learning framework based on counterfactual reasoning for medical VLMs. It comprises \textbf{BaseAnatomy}, a geometry-supervised representation learning module, and \textbf{$Δ$Phenotype}, a counterfactual incremental representation learning module. BaseAnatomy provides fine-grained geometric supervision through spatial relationships across and within anatomical structures. $Δ$Phenotype computes the representation increment between lesion representations and their corresponding normal anatomical representations, and supervises increments associated with the same phenotype to cluster in the representation space. Experiments on \textit{ReXGroundingCT} and \textit{LIDC-IDRI} demonstrate that $Δ$Representation effectively structures pathological phenotype representations and improves lesion grounding and phenotype characterization accuracy in medical VLMs. Code is available at https://anonymous.4open.science/r/deltarep-CF6D.