arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于自动化病理诊断与报告生成的视觉语言模型基准测试

Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

Yumi Lee, Harim Oh, Hyoryung Kim, Minji Kim, Eunsu Kim, Hyeseong Lee, Junya Fukuoka, Andrey Bychkov, Jijgee Munkhdelger, Rajiv Kumar Kaushal, Ayushi Sahay, Rajni Yadav, Bharathi Prabakaran, Sulen Sarioglu, Serdar Balcı, Ilknur Turkmen, Yuri Tolkach, Christian Harder, Julian Westerdorf, Reinhard Buettner, Audun Ljone Henriksen, Sepp De Raedt, Byung Hyun Lee, Sungjin Lim, Joohoon Lee, Gwanghyun Kim, Se Young Chun, Suryakant Singh, Saarthak Kapse, Prateek Prasanna, Kyung A Kim, Yousun Kang, Sehwan Yoo, Sungman Hong, Shubham Innani, Michael Feldman, Spyridon Bakas, Ujjwal Baid, Prasad Dutande, Suhas Gajare, Bhakti Baheti, Serkan Sökmen, Ece Tuğba Cebeci, Ahmet Halıcı, Musa Balcı, Kardelen Peçenek, Srividhya Sainath, Kyongseok Jang, Messi H. J. Lee, Noorul Wahab, Bodong Du, Jiaming Zhang, Qixiang Zhang, Jang-Hwan Choi, Sangjeong Ahn

arXiv 2609.00866首次发表:更新:

发表机构

Ewha Womans University; Korea University Anam Hospital; Korea University College of Medicine; Memorial Health Group; Kameda Medical Center; Nagasaki University Graduate School of Biomedical Sciences; Tata Memorial Hospital; All India Institute Of Medical Sciences Delhi; University Hospital Cologne, Medical Faculty, University of Cologne; Institute for Cancer Genetics and Informatics; Seoul National University; Stony Brook University; Yonsei University College of Medicine; Tokyo Polytechnic University; Nanyang Technological University; Korea University; Indiana University School of Medicine; Emory University; Shri Guru Gobind Singhji Institute of Engineering and Technology; Viseur AI; EKFZ TU Dresden (KatherLab); MTS Company; University of Warwick; The Hong Kong University of Science and Technology; Harbin Institute of Technology(梨花女子大学; 高丽大学安岩医院; 高丽大学医学院; 纪念健康集团; 龟田医疗中心; 长崎大学生物医学科学研究生院; 塔塔纪念医院; 全印医学科学研究所德里分院; 科隆大学医院、科隆大学医学院; 癌症遗传学与信息学研究所; 首尔大学; 石溪大学; 延世大学医学院; 东京工艺大学; 南洋理工大学; 高丽大学; 印第安纳大学医学院; 埃默里大学; 斯里古鲁戈宾德辛格吉工程技术学院; 维瑟人工智能公司; 德累斯顿工业大学EKFZ(凯瑟实验室); MTS公司; 华威大学; 香港科技大学; 哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究建立了含约10500对样本的泛亚WSI-报告数据集及REG 2025基准,评估多模态病理模型,发现需结构化表示等提升性能,指出数字幻觉等局限,为相关模型设计提供依据。

AI 中文摘要

视觉语言模型(VLMs)的快速推进加速了计算病理学的发展;然而,基于全切片图像(WSI)的病理报告生成仍受限于大规模WSI-报告数据集的稀缺,以及将空间分布的视觉模式映射为结构化临床文本的复杂性。为解决这一问题,我们引入了由五家机构整理的临床 curated泛亚WSI-报告数据集,包含约10500对样本,并通过MICCAI挑战赛建立了REG 2025基准,用于系统评估多模态模型。我们分析了提交的方法,涵盖预训练VLMs、多实例学习框架、分层专家模型、检索增强生成及跨模态Transformer。结果表明,仅使用VLM不足以实现优异性能,表现最佳的方法得益于结构化报告表示、分层诊断分解及有效的多模态 grounding。我们还发现了关键局限,包括定量属性估计的不稳定性(如数字幻觉)和诊断过度细化的倾向,部分错误类似常规病理学中的已知诊断陷阱。这些发现确立了REG 2025作为评估基于WSI的结构化报告生成及计算病理学中视觉语言理解的基准,为设计临床 grounded的多模态病理模型提供了见解。

英文摘要

The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the complexity of mapping spatially distributed visual patterns to structured clinical text. To address this, we introduce a clinically curated Pan-Asia WSI--report dataset of approximately 10,500 pairs from five institutions and establish the REG 2025 benchmark through a MICCAI challenge for systematic evaluation of multimodal models. We analyze submitted methods spanning pretrained VLMs, multiple-instance learning frameworks, hierarchical expert models, retrieval-augmented generation, and cross-modal Transformers. Rather than indicating that VLM use alone was sufficient for superior performance, the results suggest that top-performing methods benefited from structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding. We identify key limitations, including instability in quantitative attribute estimation (e.g., numeric hallucination) and a tendency toward diagnostic overspecification, with some errors resembling known diagnostic pitfalls in routine pathology. These findings establish REG 2025 as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑