arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32352cs.CVcs.AI

EyeVQA:从识别到空间定位的眼科视觉语言模型基准测试

EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding

Gujie Shao, Zixun Xie, Xuechun Xing, Ruixiang Wang, Ziyun Lan, Yanlin Qi, Gangyi Zhang, Yuxin Yang, Dawei Li, Haiming Tang

首次发表
浏览论文内容

中文总结 AI 辅助

EyeVQA是一个统一眼科VQA基准,含20,000个问答对,覆盖六类疾病和七种问题类型,44.5%需跨图像推理;基准测试14个VLM,最佳得分仅62.8,揭示空间定位和跨任务泛化的显著不足。

中文摘要 AI 辅助

视觉语言模型(VLMs)在医学图像理解方面展现出日益增长的潜力,但其在眼科影像中的能力仍未得到充分表征。现有的眼科数据集通常针对单一疾病或专门任务设计,难以系统评估VLMs能否超越疾病识别,迈向比较推理和细粒度空间定位。我们提出EyeVQA,一个统一的视觉问答基准,用于全面评估眼科VLM。EyeVQA由21个可用的眼科数据集构建,包含20,000个问答对,涵盖六类疾病和七种问题类型:单选、多选、变量选择、判断、排序、点定位和边界框。金标准答案由源提供的诊断、严重程度分级、临床发现、分割掩膜、边界框和解剖标志确定性推导,无需依赖模型生成的注释即可实现可复现评估。值得注意的是,44.5%的问题需要跨多张图像进行推理,将评估扩展到传统单图像医学VQA之外。我们在统一的零样本协议下对十四个代表性通用、科学和医学专用VLM进行基准测试。表现最佳的模型仅达到62.8的总体得分,而在空间定位和跨任务泛化方面仍存在显著差距。这些结果凸显了当前VLM在全面眼科视觉理解方面的局限性,并将EyeVQA确立为开发更可靠、更具空间定位能力的眼科多模态模型的诊断基准。项目页面可从此https URL获取。

英文摘要

Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or specialized tasks, making it difficult to systematically evaluate whether VLMs can move beyond disease recognition toward comparative reasoning and fine-grained spatial grounding. We introduce EyeVQA, a unified visual question answering benchmark for comprehensive evaluation of ophthalmic VLMs. EyeVQA is constructed from 21 available ophthalmic datasets and contains 20,000 question-answer pairs spanning six disease groups and seven question types: Single-Choice, Multi-Select, Variable-Select, True-False, Ranking, Point Location, and Bounding Box. Gold answers are deterministically derived from source-provided diagnoses, severity grades, clinical findings, segmentation masks, bounding boxes, and anatomical landmarks, enabling reproducible evaluation without relying on model-generated annotations. Notably, 44.5% of the questions require reasoning across multiple images, extending evaluation beyond conventional single-image medical VQA. We benchmark fourteen representative general-purpose, scientific, and medically specialized VLMs under a unified zero-shot protocol. The best-performing model only achieves an overall score of 62.8, while substantial gaps remain in spatial grounding and cross-task generalization. These results highlight the limitations of current VLMs in comprehensive ophthalmic visual understanding and establish EyeVQA as a diagnostic benchmark for developing more reliable and spatially grounded ophthalmic multimodal models. The project page is available at https://github.com/PKUTHM/EyeVQA.

补充信息

↑