AI 中文总结
研究病理学视觉语言模型评估中被忽视的问题,提出含多类型样本的PathBind基准,通过对多个VLMs评估发现,当前病理学VLMs在答案性能与视觉语义绑定间有显著差距。
AI 中文摘要
病理学视觉语言模型(VLMs)近来发展迅速,通常通过病理学VQA基准上的答案准确性进行评估。然而,深入研究当前评估发现三个被忽视的问题:视觉证据并非总是必要;领域训练能提高准确性,但视觉绑定方面没有成比例提升;实体级注意力分散且与查询特定性弱相关。这些问题会导致对病理学VLMs实际多模态能力的重大误判。为此,提出包含2600个样本的PathBind基准,包括PathBind-VQA、PathBind-PTA和PathBind-Grounding。对18个代表性VLMs在PathBind的VQA样本及五个现有病理学VQA基准上进行评估,还对10个VLMs在PathBind-Grounding和PathVG上评估。结果表明当前病理学VLMs在答案性能和视觉语义绑定之间仍存在巨大差距。
英文摘要
Pathology vision-language models (VLMs) have recently progressed rapidly and are commonly evaluated by answer accuracy on pathology VQA benchmarks. However, we dig into current evaluations and identify three overlooked issues: 1) Visual evidence is not always necessary. For instance, Gemini-3-Pro achieves 53.5% average accuracy across 5 VQA benchmarks without any visual input. 2) Domain training can improve accuracy without proportional gains in visual binding. Compared with Qwen2.5-VL-7B, Patho-R1-7B exhibits a 5.8-point lower multimodal gain and a 3.7-point lower attention IoU. 3) Entity-level attention is diffuse and weakly query-specific. On PathVG, attention maps remain highly correlated across different entity queries. These issues can lead to substantial misjudgments of pathology VLMs' actual multimodal capabilities. To this end, we present PathBind, a benchmark comprising 2,600 samples: PathBind-VQA with 1,500 questions across six dimensions, PathBind-PTA with 600 questions from a private pathology teaching atlas, and PathBind-Grounding with 500 expert-curated region-level samples. Each component undergoes task-specific automated filtering and expert review to reduce textual shortcuts and improve entity-region correspondence. We evaluate 18 representative VLMs on VQA samples of PathBind and five existing pathology VQA benchmarks, and further evaluate 10 VLMs on PathBind-Grounding and PathVG. Results show that current pathology VLMs still exhibit a substantial gap between answer-side performance and visual-semantic binding.