arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PathView-Bench:多模态大语言模型能否实现病理图像的细粒度多尺度理解?

PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?

Zongyi Chen, Yu Liang, Jie Lin, Liansheng Wang

arXiv 2607.28318首次发表:更新:

发表机构

National Institute for Data Science in Health and Medicine, Xiamen University(厦门大学健康与医学数据科学国家研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有病理多模态基准的不足,推出PathVU基准,评估发现主流MLLMs在多尺度病理图像细粒度视觉任务上存在显著局限,为相关模型的开发评估提供了可复现基础。

AI 中文摘要

多模态大语言模型(MLLMs)正越来越多地被用于分析病理图像。然而,主流的病理多模态基准主要对最终诊断答案、图像描述或病理报告进行评分,这些评估难以深入了解模型是否掌握了病理推理与决策所需的多尺度视觉内容。我们推出PathVU,这是一个针对计算病理学中细粒度多尺度视觉理解的视觉锚定基准。该基准基于23个带有人工监督标签和空间注释的公开病理成像数据集构建,从Region FOV(高分辨率局部区域)和Slide FOV(宏观全切片视图)两个视场维度评估MLLM的理解能力。通过将原始注释转换为确定性任务目标,PathVU支持对区域定位、视觉识别、数量估计、空间推理以及上下文不足判断进行程序化评分。该基准包含14个VQA风格任务、61673张图像、308070个样本,覆盖28个器官,包含7253526条注释。我们对18个具有代表性的通用型、医学领域型及病理专用型MLLMs进行评估后发现,即使是先进模型,在多尺度病理图像的细粒度视觉任务上也存在显著局限。PathVU为开发和评估具备明确多尺度视觉理解能力的病理MLLMs提供了可复现的基础。

英文摘要

Multimodal large language models (MLLMs) are increasingly used to analyze pathology images. However, dominant multimodal benchmarks in pathology mainly score final diagnostic answers, captions, or reports. These evaluations provide limited insight into whether a model understands the multiscale visual content needed for pathology reasoning and decision-making. We introduce PathVU, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology. Built from 23 public pathology imaging datasets with human-supervised labels and spatial annotations, PathVU evaluates MLLM understanding in two fields of view: Region FOV for high-resolution local regions and Slide FOV for macro whole-slide views. By converting raw annotations into deterministic task targets, PathVU enables programmatic scoring of region localization, visual recognition, quantity estimation, spatial reasoning, and insufficient-context judgment. The benchmark contains 14 VQA-style tasks, 61,673 images, and 308,070 samples across 28 organs and 7,253,526 annotations. Evaluating 18 representative general-purpose, medical-domain, and pathology-oriented MLLMs, we observe substantial limitations even in advanced models on fine-grained visual tasks across multiscale pathology images. PathVU provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑