GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards
GI-Bench: 一个全景基准测试,揭示多模态大语言模型在内窥镜检查中与临床标准相比的知识-经验脱节
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
AI总结 GI-Bench通过评估多模态大语言模型在胃肠内窥镜中的性能,揭示其在诊断推理和空间定位方面的不足,发现模型在语言流畅度上优于人类,但事实准确性较低。
Comments 45 pages, 17 figures, 6 tables. Leaderboard available at: https://roterdl.github.io/GIBench/ . Includes supplementary material