EH-Benchmark:眼科幻觉基准与智能体驱动的自顶向下可追溯推理工作流
EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
- Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR)(高性能计算研究所,科技研究局(A*STAR))
- Centre for Innovation and Precision Eye Health(创新与精准眼健康中心)
- Department of Ophthalmology, NUHS Tower Block, Level 7, 1E Kent Ridge Road, Singapore, 119228(眼科部,NUHS塔楼7层,1E Kent Ridge Road,新加坡,119228)
- Singapore Eye Research Institute, Singapore National Eye Centre, 20 College Road, Singapore, 169856(新加坡眼研究 institute,新加坡国家眼科中心,20 College Road,新加坡,169856)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对医学大语言模型在眼科诊断中的幻觉问题,本研究提出EH-Benchmark眼科幻觉评估基准,将幻觉分为视觉理解与逻辑组合两大类,并设计三阶段多智能体推理框架,有效缓解幻觉、提升诊断准确性与可靠性。
AI中文摘要:
医学大语言模型(MLLMs)在眼科诊断中发挥着关键作用,在应对威胁视力的疾病方面具有巨大潜力。然而,其准确性受到幻觉问题的限制,这些幻觉源于眼科知识有限、视觉定位与推理能力不足,以及多模态眼科数据稀缺,这些因素共同阻碍了精准的病变检测与疾病诊断。此外,现有医学基准无法有效评估各类幻觉,也无法提供可落地的缓解方案。为应对上述挑战,我们提出了EH-Benchmark,这是一款专为评估MLLMs幻觉问题设计的新型眼科基准。我们根据具体任务和错误类型,将MLLMs的幻觉分为两大类:视觉理解类与逻辑组合类,每类包含多个子类。鉴于MLLMs主要依赖基于语言的推理而非视觉处理,我们提出了一个以智能体为核心的三阶段框架,包括知识层检索阶段、任务层案例研究阶段和结果层验证阶段。实验结果表明,我们的多智能体框架能显著缓解两类幻觉问题,提升准确性、可解释性与可靠性。我们的项目可在https://github.com/ppxy1/EH-Benchmark获取。
英文摘要:
Medical Large Language Models (MLLMs) play a crucial role in ophthalmic diagnosis, holding significant potential to address vision-threatening diseases. However, their accuracy is constrained by hallucinations stemming from limited ophthalmic knowledge, insufficient visual localization and reasoning capabilities, and a scarcity of multimodal ophthalmic data, which collectively impede precise lesion detection and disease diagnosis. Furthermore, existing medical benchmarks fail to effectively evaluate various types of hallucinations or provide actionable solutions to mitigate them. To address the above challenges, we introduce EH-Benchmark, a novel ophthalmology benchmark designed to evaluate hallucinations in MLLMs. We categorize MLLMs' hallucinations based on specific tasks and error types into two primary classes: Visual Understanding and Logical Composition, each comprising multiple subclasses. Given that MLLMs predominantly rely on language-based reasoning rather than visual processing, we propose an agent-centric, three-phase framework, including the Knowledge-Level Retrieval stage, the Task-Level Case Studies stage, and the Result-Level Validation stage. Experimental results show that our multi-agent framework significantly mitigates both types of hallucinations, enhancing accuracy, interpretability, and reliability. Our project is available at https://github.com/ppxy1/EH-Benchmark.