AI 中文总结
该研究构建多模态医学视觉定位数据集LocAnyMed-200K及理由增强子集LocAnyMed-CoT-20K,通过微调LocateAnything-3B提升医学定位性能,为相关研究提供统一基础。
AI 中文摘要
医学视觉定位将自由格式的临床查询与医学图像中的空间证据关联起来,是可解释医学人工智能的重要组成部分。然而,通用定位模型主要在自然图像上训练,而现有的医学定位资源在成像模态、数据集和任务形式上仍呈碎片化。为解决这一差距,我们构建了LocAnyMed-200K,这是一个多模态医学视觉定位数据集,包含约20万张图像-查询-答案示例,涵盖计算机断层扫描、光学医学成像、超声和X射线。我们将异构检测与定位资源统一为一种统一的自由格式指令格式,支持一个或多个边界框、点坐标,以及针对负查询的无目标输出。在LocAnyMed-200K上对LocateAnything-3B进行全参数微调,使保留的评估分割上的F1@IoU 0.50从10.64提升至85.59,证明大规模特定领域监督可使通用定位模型具备有效的医学定位能力。除空间坐标外,临床可解释的定位系统还应传达支撑其预测的证据。因此,我们衍生了LocAnyMed-CoT-20K,这是一个 rationale(理由)增强子集,通过结构化推理关联解剖学背景、视觉观察和空间结论,并通过微调进一步提升跨源泛化能力。这些资源共同为研究异构医学成像模态下的定位准确性和理由质量提供了统一基础。代码可在此https URL获取。
英文摘要
Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medical artificial intelligence. However, general-purpose grounding models are predominantly trained on natural images, while existing medical localization resources remain fragmented across imaging modalities, datasets, and task formulations. To address this gap, we construct LocAnyMed-200K, a multimodal medical visual grounding dataset containing approximately 200K image-query-answer examples across computed tomography, optical medical imaging, ultrasound, and X-ray. We harmonize heterogeneous detection and localization resources into a unified free-form instruction format that supports one or multiple bounding boxes, point coordinates, and no-target outputs for negative queries. Full-parameter fine-tuning of LocateAnything-3B on LocAnyMed-200K improves F1@IoU 0.50 from 10.64 to 85.59 on a held-out evaluation split, demonstrating that large-scale domain-specific supervision can equip a general grounding model with effective medical localization capabilities. Beyond spatial coordinates, a clinically interpretable grounding system should also communicate the evidence supporting its prediction. We therefore derive LocAnyMed-CoT-20K, a rationale-augmented subset that connects anatomical context, visual observations, and spatial conclusions through structured reasoning and further improves cross-source generalization through fine-tuning. Together, these resources provide a unified foundation for studying both localization accuracy and rationale quality across heterogeneous medical imaging modalities. The code is publicly available at https://github.com/MiliLab/LocAnyMed.
CommentsTechnical report; work in progress. 28 pages, 5 figures, and 16 tables. Code: https://github.com/MiliLab/LocAnyMed