发表机构
University of International Relations(国际关系学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在增强医学多模态大语言模型的 3D 空间推理。通过新型逐片数据合成范式构建结构化推理数据集,用其对二维预训练模型进行指令微调,提升了模型在 3D 医学基准测试中的性能,缩小了与原生 3D 架构的差距。
AI 中文摘要
多模态大语言模型在二维医学图像理解方面取得了显著成功,但其向三维容积成像的扩展仍受标注成本高昂和数据集不透明的阻碍。当前数据格式通常无法捕捉明确的临床推理。为此,我们引入了一个通过新型逐片数据合成范式构建的大规模结构化推理数据集。该范式受放射科医生真实诊断流程启发,通过分解复杂的三维阅读过程来模拟视觉认知,将全局临床先验转化为细粒度的逐片观察,进而合成可解释的思维链。合成推理框架强化了关键临床原则。我们用合成数据对标准二维预训练的多模态大语言模型基线进行指令微调以增强其容积理解能力。多项三维医学基准测试表明我们的方法比二维基线有显著性能提升,缩小了与资源密集型原生三维架构的性能差距,且无需计算昂贵的三维特定预训练。完整存储库公开可用。
英文摘要
While Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in 2D medical image understanding, their extension to 3D volumetric imaging remains hindered by prohibitive annotation costs and dataset opacity. Current data formats, predominantly consisting of rigid Visual Question Answering (VQA) pairs or unstructured final clinical reports, typically fail to capture explicit clinical reasoning. To address this limitation, we introduce a large-scale structured reasoning dataset constructed via a novel slice-wise data synthesis paradigm. Inspired by the genuine diagnostic workflow of radiologists, this paradigm models visual cognition by decomposing the complex 3D reading process, translating global clinical priors into fine-grained, per-slice observations that are subsequently synthesized into an interpretable Chain-of-Thought (CoT). Crucially, this synthesized reasoning framework enforces essential clinical principles: sequential spatial tracking, multi-slice spatial awareness for artifact mitigation, and differential exclusion. To validate this approach, we instruction-tune a standard 2D-pretrained MLLM baseline using the synthesized data to enhance its volumetric comprehension. Comprehensive evaluations across multiple 3D medical benchmarks demonstrate that our method yields significant performance improvements over the 2D baseline. Furthermore, the resulting model exhibits robust spatial reasoning capabilities and rivals resource-intensive native 3D architectures, effectively bridging the performance gap. Ultimately, this data-centric strategy unlocks deep volumetric understanding and highly interpretable clinical logic without requiring computationally expensive 3D-specific pre-training. The complete repository, including datasets and training workflows, is publicly available at https://github.com/2020420145009/hounsfield.