arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

病理学视觉语言模型真的能‘看到’病理学吗?

Do Pathology Vision-Language Models Truly See Pathology?

Chengyang Zhang, Wenchuan Zhang, Bo Li, Xinyu Liu, Jiaming Yang, Mengran Li, Chenxun Deng, Jie Chen, Yang Zhang, Wei Ju, Yuhao Yi, Hong Bu, Jiancheng Lv

arXiv 2607.21065首次发表:更新:

AI 中文总结

研究病理学视觉语言模型评估中被忽视的问题,提出含多类型样本的PathBind基准,通过对多个VLMs评估发现,当前病理学VLMs在答案性能与视觉语义绑定间有显著差距。

AI 中文摘要

病理学视觉语言模型(VLMs)近来发展迅速,通常通过病理学VQA基准上的答案准确性进行评估。然而,深入研究当前评估发现三个被忽视的问题:视觉证据并非总是必要;领域训练能提高准确性,但视觉绑定方面没有成比例提升;实体级注意力分散且与查询特定性弱相关。这些问题会导致对病理学VLMs实际多模态能力的重大误判。为此,提出包含2600个样本的PathBind基准,包括PathBind-VQA、PathBind-PTA和PathBind-Grounding。对18个代表性VLMs在PathBind的VQA样本及五个现有病理学VQA基准上进行评估,还对10个VLMs在PathBind-Grounding和PathVG上评估。结果表明当前病理学VLMs在答案性能和视觉语义绑定之间仍存在巨大差距。

英文摘要

Pathology vision-language models (VLMs) have recently progressed rapidly and are commonly evaluated by answer accuracy on pathology VQA benchmarks. However, we dig into current evaluations and identify three overlooked issues: 1) Visual evidence is not always necessary. For instance, Gemini-3-Pro achieves 53.5% average accuracy across 5 VQA benchmarks without any visual input. 2) Domain training can improve accuracy without proportional gains in visual binding. Compared with Qwen2.5-VL-7B, Patho-R1-7B exhibits a 5.8-point lower multimodal gain and a 3.7-point lower attention IoU. 3) Entity-level attention is diffuse and weakly query-specific. On PathVG, attention maps remain highly correlated across different entity queries. These issues can lead to substantial misjudgments of pathology VLMs' actual multimodal capabilities. To this end, we present PathBind, a benchmark comprising 2,600 samples: PathBind-VQA with 1,500 questions across six dimensions, PathBind-PTA with 600 questions from a private pathology teaching atlas, and PathBind-Grounding with 500 expert-curated region-level samples. Each component undergoes task-specific automated filtering and expert review to reduce textual shortcuts and improve entity-region correspondence. We evaluate 18 representative VLMs on VQA samples of PathBind and five existing pathology VQA benchmarks, and further evaluate 10 VLMs on PathBind-Grounding and PathVG. Results show that current pathology VLMs still exhibit a substantial gap between answer-side performance and visual-semantic binding.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑