AI 中文总结
研究探讨多模态大语言模型能否理解OCT,引入OCT - Bench基准,包含多维度细粒度任务,基于多个数据集构建大量选择题,评估20个代表性模型,发现当前模型理解OCT能力不足,为评估模型和推动OCT理解提供基础。
AI 中文摘要
光学相干断层扫描(OCT)成像对视网膜疾病的诊断和治疗至关重要。尽管多模态大语言模型(MLLMs)在医学图像分析中显示出巨大潜力,但现有基准大多将OCT理解简化为粗粒度疾病分类或孤立的视觉问答,完整认知过程评估不足。为此引入OCT - Bench,一个致力于OCT图像理解的综合基准。它包含10,076个高质量选择题,基于七个公共数据集的4,137张OCT图像构建。按照临床解释工作流程,建立了由20个细粒度任务组成的分层能力分类法,涵盖多个方面。系统评估20个代表性MLLMs,结果表明当前模型在可靠理解OCT方面仍有很大差距,医学领域适应和模型规模增加也未持续提升性能。OCT - Bench为全面细粒度评估MLLMs提供基础,有助于识别能力瓶颈并推动基于临床的OCT理解。
英文摘要
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained disease classification or isolated visual question answering, leaving the complete cognitive process from visual perception to clinical reasoning insufficiently evaluated. To address this limitation, we introduce OCT-Bench, a comprehensive benchmark dedicated to OCT image understanding. OCT-Bench comprises 10,076 high-quality multiple-choice questions constructed from 4,137 OCT images across seven public datasets. Following the real-world clinical interpretation workflow, we establish a hierarchical capability taxonomy consisting of 20 fine-grained tasks across three dimensions: Perception, Cognition, and Reasoning. These tasks cover a broad range of capabilities, including imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, therapeutic decision-making, and prognostic management. We systematically evaluate 20 representative MLLMs, including proprietary models, open-source general-purpose models, and medical-domain models. Experimental results demonstrate that current models remain substantially short of reliable OCT understanding. Moreover, neither medical-domain adaptation nor increased model scale consistently improves performance across capability levels. OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.