arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PathAgentBench:在全切片病理图像上对寻求证据的视觉语言模型进行基准测试

EviPathBench: Benchmarking Evidence Acquisition and Reasoning in Vision-Language Models for Whole-Slide Pathology

Dankai Liao, Tianyi Zhang, Yufeng Wu, Xinyue Zhang, Qiaochu Xue, Zeyu Liu, Dachun Zhao, Linghan Cai, Yueming Jin

arXiv 2607.19261首次发表:更新:

发表机构

National University of Singapore; PuzzleLogic Pte Ltd; Harbin Institute of Technology, Shenzhen; Peking Union Medical College Hospital(新加坡国立大学; 拼图逻辑私人有限公司; 哈尔滨工业大学深圳校区; 北京协和医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对全切片病理图像诊断中证据获取与整合问题,引入PathAgentBench基准,评估视觉语言模型的四项互补能力,包含多模型评估结果,揭示了证据推理与获取间的差距,为改进病理模型提供统一框架。

AI 中文摘要

全切片图像(WSI)诊断需要识别诊断相关区域,跨放大倍数检查并整合多尺度证据。但现有病理基准大多在预裁剪补丁或预提取的切片特征上评估模型,其从千兆像素WSI直接获取证据的能力未经充分测试。我们引入PathAgentBench,用于评估寻求证据的视觉语言模型的四项互补能力:证据解释的图像到文本匹配、证据验证的文本到图像检索、证据获取的诊断区域定位以及证据整合的多尺度推理。该基准以诊断树形式组织,包含1822张TCGA WSI和17135条由十位认证病理学家注释的诊断路径。使用190张带详细注释的乳腺癌WSI私人队列评估自主全切片探索。我们评估了20个通用、医学和病理专用模型。领先的开放权重模型在多尺度推理中准确率超过93%,在两个跨模态匹配任务中准确率超过50%。相比之下,诊断区域定位仍具挑战性:最佳文本引导平均交并比低于0.09,不如简单的基于中心的启发式方法。在自主探索中,无条件命中率从低倍放大时的0.522降至中倍放大时的0.185和高倍放大时的0.020。这些结果揭示了在整理好的证据上推理与直接从WSI获取证据之间的明显差距。PathAgentBench为测量和改进寻求证据的病理模型提供了统一框架。

英文摘要

Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence. However, most pathology benchmarks evaluate models on pre-cropped patches or pre-extracted slide features, leaving their ability to acquire evidence from gigapixel WSIs largely untested. We introduce EviPathBench, a benchmark for evaluating evidence acquisition and reasoning in vision-language models (VLMs) for whole-slide pathology. It evaluates four capabilities: image-to-text matching for evidence interpretation, text-to-image retrieval for evidence verification, diagnostic-region localization for evidence acquisition, and multi-scale reasoning for evidence integration. The benchmark is organized as a diagnostic tree linking nested regions across magnifications with scale-specific findings and path-level diagnoses. It contains 1,822 TCGA WSIs and 17,135 diagnostic paths annotated by ten board-certified pathologists. A private cohort of 190 breast cancer WSIs with detailed annotations further evaluates autonomous whole-slide exploration. We evaluate 19 VLMs spanning general-purpose, medical, and pathology-specialized families, plus one text-only reference model. Leading open-weight models achieve over 93% accuracy in multi-scale reasoning and over 50% in both cross-modal matching tasks. In contrast, diagnostic-region localization remains challenging: the best text-guided mean intersection-over-union is below 0.09, underperforming a center-based heuristic. During autonomous exploration, the unconditional hit rate drops from 0.522 at low magnification to 0.185 at intermediate magnification and 0.020 at high magnification. These results reveal a pronounced gap between reasoning over curated evidence and acquiring it from WSIs. EviPathBench provides a unified framework for measuring and improving both capabilities.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑