医疗多模态大语言模型的视点临床意图理解能力基准测试
Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
- Shenzhen University(深圳大学)
- Stanford University(斯坦福大学)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出MedGaze-Bench,首个评估医疗多模态大语言模型视点临床意图理解能力的基准测试,通过三维意图框架和陷阱QA机制,揭示现有模型在手术、急救和诊断任务中对意图理解的不足。
AI中文摘要:
医疗多模态大语言模型(Med-MLLMs)需要视点临床意图理解能力以实现实际应用,但现有基准测试未能评估这一关键能力。为解决这些挑战,我们引入了MedGaze-Bench,这是首个利用临床人员目光作为认知光标来评估手术、急救模拟和诊断解释中意图理解的基准测试。我们的基准测试解决了三个根本性挑战:解剖结构的视觉同质性、临床工作流程中的严格时间因果依赖性以及隐含的安全规程遵守。我们提出了一种三维临床意图框架,评估:(1)空间意图:在视觉噪声中区分精确目标,(2)时间意图:通过回顾性和前瞻性推理推断因果逻辑,(3)标准意图:通过安全检查验证规程合规性。除了准确性指标外,我们引入了陷阱QA机制,通过惩罚幻觉和认知趋同来压力测试临床可靠性。实验发现当前MLLMs在视点意图上存在问题,主要是过度依赖全局特征,导致生成虚假观察和无批判性接受无效指令。
英文摘要:
Medical Multimodal Large Language Models (Med-MLLMs) require egocentric clinical intent understanding for real-world deployment, yet existing benchmarks fail to evaluate this critical capability. To address these challenges, we introduce MedGaze-Bench, the first benchmark leveraging clinician gaze as a Cognitive Cursor to assess intent understanding across surgery, emergency simulation, and diagnostic interpretation. Our benchmark addresses three fundamental challenges: visual homogeneity of anatomical structures, strict temporal-causal dependencies in clinical workflows, and implicit adherence to safety protocols. We propose a Three-Dimensional Clinical Intent Framework evaluating: (1) Spatial Intent: discriminating precise targets amid visual noise, (2) Temporal Intent: inferring causal rationale through retrospective and prospective reasoning, and (3) Standard Intent: verifying protocol compliance through safety checks. Beyond accuracy metrics, we introduce Trap QA mechanisms to stress-test clinical reliability by penalizing hallucinations and cognitive sycophancy. Experiments reveal current MLLMs struggle with egocentric intent due to over-reliance on global features, leading to fabricated observations and uncritical acceptance of invalid instructions.