From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
从复合图形到综合理解:开发一种从生物医学文献中基于医疗多图像基准测试和验证的多模态大语言模型
机构 * Department of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven, CT 06510, USA(耶鲁医学院生物医学信息学与数据科学系,耶鲁大学,新 Haven,CT 06510,USA) ; School of Medicine, University of Puerto Rico, San Juan, PR 00921, USA(波多黎各大学医学院,波多黎各,San Juan,PR 00921,USA)
专题命中 生物医学文本 :biomedical(title,abstract);diagnosis(abstract);分类 cs.CV
AI总结 本文提出M3LLM,一种基于生物医学文献中复合图像的多模态大语言模型,通过分而治之策略提升多图像理解能力,实验证明其在多图像、单图像等场景中表现优异。