医学视觉模型是否对解剖结构进行推理?探究学习到的视觉表示的空间归纳偏置
Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations
AI总结:
本研究构建SPAR-Bench探测任务,发现医学视觉模型仅记忆标准解剖结构,缺乏特定患者图像内结构对比的推理能力,池化探测会低估其表示能力。
AI中文摘要:
解读CT扫描图像需要对比两侧结构、判断器官间距及知晓各器官位置。医学视觉编码器通常通过诊断准确率或集成多模态系统评估,而后者的故障难以归因,因此目前尚不清楚这些编码器的表示是否支持上述空间推理能力。我们构建了SPAR-Bench,这是针对多器官腹部CT的八项探测任务,涵盖坐标定位、关系推理和空间查询三类任务,并将其应用于五种架构配置和三种医学基础模型(含冻结模型与微调模型)。在同一切片内要求对比的探测任务表现仅达随机水平,且预训练规模、微调或架构均未缩小差距;在特定领域看似已解决的探测任务,在零样本迁移时降至随机水平,表明其准确率反映的是对标准解剖结构的记忆而非对图像的计算。使用池化头而非全部标记读取相同的冻结特征,可使关系恢复率从0.7%提升至67.8%,说明池化探测低估了表示所蕴含的能力。编码器能良好回答的问题,四个开放权重多模态大语言模型(MLLMs)回答时也仅达随机水平。我们的结果表明,这些编码器携带的是器官通常位置的图谱,而非特定患者图像内结构对比的机制。代码和数据将在此https URL发布。
英文摘要:
Interpreting a CT scan means comparing structures on either side, judging how far apart organs sit, and knowing where each one belongs. Medical vision encoders are evaluated on diagnostic accuracy, or through assembled multimodal systems where a failure is hard to attribute, so it remains unclear whether their representations support any of this. We construct SPAR-Bench, eight probes over multi-organ abdominal CT that separate coordinate localization, relational reasoning, and spatial queries, and apply them to five architectural configurations and three medical foundation models, frozen and finetuned. Probes that ask for a comparison within the slice stay at chance, and neither pretraining scale, finetuning, nor architecture closes the gap. Probes that appear solved in domain fall to chance under zero-shot transfer, indicating that their accuracy reflects recall of canonical anatomy rather than computation over the image. Reading the same frozen features with a pooled head rather than the full set of tokens moves relational recovery from 0.7% to 67.8%, so pooled probing understates what a representation holds. Questions the encoders answer well are answered at chance by four open-weight MLLMs. Our results suggest these encoders carry a map of where organs usually lie, and little of the machinery for comparing structures within a particular patient. Code and data will be available at https://spar-bench.github.io.