发表机构
The University of British Columbia; Vector Institute(不列颠哥伦比亚大学; 矢量研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出基于位置感知细粒度表示学习的医学视觉基础模型LoFi,构建大规模医学定位数据集MedG,在多项医学视觉任务中性能优于现有模型。
AI 中文摘要
细粒度视觉表示对于医学图像分析至关重要,尤其是当诊断相关证据细微且空间定位明确时。现代基于Transformer的医学视觉编码器因此必须学习既具有临床意义又空间一致的 patch 级表示。若缺少这些特性,大型视觉语言模型(LVLMs)将建立在模糊的视觉基础上,限制其生成临床可靠且空间定位明确的响应的能力。然而,现有的医学视觉编码器训练策略很少能同时实现这两个目标。图像-文本对齐主要在图像级别提供临床意义监督,而诊断证据的空间定位约束较弱。相比之下,自监督学习可促进空间一致性,但缺乏区分视觉相似但临床不同区域所需的语义监督。为解决这一差距,我们提出LoFi,一种基于位置感知细粒度表示学习的医学视觉基础模型。LoFi在定位和接地字幕目标下训练视觉编码器与轻量级大型语言模型。由于这些目标需要从临床文本预测位置,反之亦然,因此无需任何显式patch级正则化即可实现空间一致性。为实现大规模训练,我们构建了MedG,这是一个包含448万张图像-文本-边界框三元组的大规模医学定位数据集,由涵盖7种模态的84个数据集整理而成。在短语定位、视觉问答和扰动下的基于区域的器官分类任务中,LoFi始终优于通用及医学视觉基础模型,以及最先进的LVLMs。代码可在this https URL获取。
英文摘要
Fine-grained visual representations are essential for medical image analysis, particularly when diagnostically relevant evidence is subtle and spatially localized. Modern transformer-based medical vision encoders must therefore learn patch-level representations that are both clinically meaningful and spatially consistent. Without these properties, large vision-language models (LVLMs) operate on an ambiguous visual foundation, limiting their ability to generate clinically reliable and spatially grounded responses. However, existing training strategies for medical vision encoders rarely achieve both objectives. Image-text alignment provides clinically meaningful supervision primarily at the image level, leaving the spatial localization of diagnostic evidence weakly constrained. In contrast, self-supervised learning promotes spatial consistency but lacks the semantic supervision needed to distinguish visually similar yet clinically distinct regions. To address this gap, we present LoFi, a medical vision foundation model built on location-aware fine-grained representation learning. LoFi trains a vision encoder with a lightweight large language model under grounding and grounded captioning objectives. Because these objectives require predicting location from clinical text and vice versa, spatial consistency emerges without any explicit patch-level regularization. To enable training at scale, we construct MedG, a large-scale medical grounding dataset of 4.48M image-text-box triplets curated from 84 datasets spanning 7 modalities. Across phrase grounding, visual question answering, and region-based organ classification under perturbations, LoFi consistently outperforms general-purpose and medical vision foundation models as well as state-of-the-art LVLMs. Code is available at https://github.com/myeongkyunkang/lofi-medg.