发表机构
Ewha Womans University; Korea University Anam Hospital; Korea University College of Medicine; Memorial Health Group; Kameda Medical Center; Nagasaki University Graduate School of Biomedical Sciences; Tata Memorial Hospital; All India Institute Of Medical Sciences Delhi; University Hospital Cologne, Medical Faculty, University of Cologne; Institute for Cancer Genetics and Informatics; Seoul National University; Stony Brook University; Yonsei University College of Medicine; Tokyo Polytechnic University; Nanyang Technological University; Korea University; Indiana University School of Medicine; Emory University; Shri Guru Gobind Singhji Institute of Engineering and Technology; Viseur AI; EKFZ TU Dresden (KatherLab); MTS Company; University of Warwick; The Hong Kong University of Science and Technology; Harbin Institute of Technology(梨花女子大学; 高丽大学安岩医院; 高丽大学医学院; 纪念健康集团; 龟田医疗中心; 长崎大学生物医学科学研究生院; 塔塔纪念医院; 全印医学科学研究所德里分院; 科隆大学医院、科隆大学医学院; 癌症遗传学与信息学研究所; 首尔大学; 石溪大学; 延世大学医学院; 东京工艺大学; 南洋理工大学; 高丽大学; 印第安纳大学医学院; 埃默里大学; 斯里古鲁戈宾德辛格吉工程技术学院; 维瑟人工智能公司; 德累斯顿工业大学EKFZ(凯瑟实验室); MTS公司; 华威大学; 香港科技大学; 哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究建立了含约10500对样本的泛亚WSI-报告数据集及REG 2025基准,评估多模态病理模型,发现需结构化表示等提升性能,指出数字幻觉等局限,为相关模型设计提供依据。
AI 中文摘要
视觉语言模型(VLMs)的快速推进加速了计算病理学的发展;然而,基于全切片图像(WSI)的病理报告生成仍受限于大规模WSI-报告数据集的稀缺,以及将空间分布的视觉模式映射为结构化临床文本的复杂性。为解决这一问题,我们引入了由五家机构整理的临床 curated泛亚WSI-报告数据集,包含约10500对样本,并通过MICCAI挑战赛建立了REG 2025基准,用于系统评估多模态模型。我们分析了提交的方法,涵盖预训练VLMs、多实例学习框架、分层专家模型、检索增强生成及跨模态Transformer。结果表明,仅使用VLM不足以实现优异性能,表现最佳的方法得益于结构化报告表示、分层诊断分解及有效的多模态 grounding。我们还发现了关键局限,包括定量属性估计的不稳定性(如数字幻觉)和诊断过度细化的倾向,部分错误类似常规病理学中的已知诊断陷阱。这些发现确立了REG 2025作为评估基于WSI的结构化报告生成及计算病理学中视觉语言理解的基准,为设计临床 grounded的多模态病理模型提供了见解。
英文摘要
The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the complexity of mapping spatially distributed visual patterns to structured clinical text. To address this, we introduce a clinically curated Pan-Asia WSI--report dataset of approximately 10,500 pairs from five institutions and establish the REG 2025 benchmark through a MICCAI challenge for systematic evaluation of multimodal models. We analyze submitted methods spanning pretrained VLMs, multiple-instance learning frameworks, hierarchical expert models, retrieval-augmented generation, and cross-modal Transformers. Rather than indicating that VLM use alone was sufficient for superior performance, the results suggest that top-performing methods benefited from structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding. We identify key limitations, including instability in quantitative attribute estimation (e.g., numeric hallucination) and a tendency toward diagnostic overspecification, with some errors resembling known diagnostic pitfalls in routine pathology. These findings establish REG 2025 as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology models.