EyePCR:面向眼科手术中细粒度感知、知识理解与临床推理的综合基准
EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery
- School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
- School of Computer Science, University of Nottingham Ningbo China(诺丁汉大学宁波校区计算机科学学院)
- Wenzhou Medical University(温州医科大学)
- School of Computer Science, University of Nottingham(诺丁汉大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究构建了眼科手术分析大规模基准EyePCR,从感知、理解、推理三维度评估MLLMs,还推出领域适配模型EyePCR-MLLM,性能比肩商业模型,为手术视频理解模型的临床可靠性提升奠基。
AI中文摘要:
多模态大语言模型(MLLMs)已展现出卓越能力,但在手术场景等高风险、特定领域场景中的性能仍未得到充分探索。为填补这一空白,我们构建了**EyePCR**——一个用于眼科手术分析的大规模基准,其以结构化临床知识为基础,从「感知」「理解」「推理」三个维度评估模型认知能力。EyePCR提供了包含超21万个视觉问答(VQA)的丰富标注语料,覆盖1048项细粒度属性以支持多视角感知评估,包含超2.5万个三元组的医学知识图谱用于理解能力评估,还设置了四项基于临床实际的推理任务。这些丰富标注可支撑深度认知分析,模拟外科医生感知视觉线索并结合领域知识做出决策的过程,从而显著提升模型的认知能力。特别地,**EyePCR-MLLM**作为Qwen2.5-VL-7B的领域适配变体,在对比模型中感知类选择题准确率最高,在理解与推理任务上优于开源模型,性能可与GPT-4.1等商业模型媲美。EyePCR揭示了现有MLLMs在手术认知方面的局限性,为手术视频理解模型的基准测试与临床可靠性提升奠定了基础。
英文摘要:
MLLMs (Multimodal Large Language Models) have showcased remarkable capabilities, but their performance in high-stakes, domain-specific scenarios like surgical settings, remains largely under-explored. To address this gap, we develop \textbf{EyePCR}, a large-scale benchmark for ophthalmic surgery analysis, grounded in structured clinical knowledge to evaluate cognition across \textit{Perception}, \textit{Comprehension} and \textit{Reasoning}. EyePCR offers a richly annotated corpus with more than 210k VQAs, which cover 1048 fine-grained attributes for multi-view perception, medical knowledge graph of more than 25k triplets for comprehension, and four clinically grounded reasoning tasks. The rich annotations facilitate in-depth cognitive analysis, simulating how surgeons perceive visual cues and combine them with domain knowledge to make decisions, thus greatly improving models' cognitive ability. In particular, \textbf{EyePCR-MLLM}, a domain-adapted variant of Qwen2.5-VL-7B, achieves the highest accuracy on MCQs for \textit{Perception} among compared models and outperforms open-source models in \textit{Comprehension} and \textit{Reasoning}, rivalling commercial models like GPT-4.1. EyePCR reveals the limitations of existing MLLMs in surgical cognition and lays the foundation for benchmarking and enhancing clinical reliability of surgical video understanding models.