arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PathLang:面向计算病理学中视觉-语言模型的以语言为中心的基准

PathLang: A Language-Centered Benchmark for Vision-Language Models in Computational Pathology

Fanqi Cheng, Kuo Gong, Shangke Liu, Beidi Zhao, Junchao Zhu, Zheyu Zhu, Leiyue Zhao, Fengbei Liu, John Cannon, Gang Wang, Zu-hua Gao, Kenji Ikemura, Yihe Yang, Yaohong Wang, Yuankai Huo, Xiaoxiao Li, Mert R. Sabuncu, Ruining Deng

arXiv 2610.11329首次发表:更新:

发表机构

New York University; Columbia University; Weill Cornell Medicine; Cornell University; University of British Columbia; Vanderbilt University; University of Pennsylvania; Johns Hopkins University; New York Medical College; BC Cancer Agency; The University of Texas MD Anderson Cancer Center; Cornell Tech(纽约大学; 哥伦比亚大学; 威尔康奈尔医学院; 康奈尔大学; 不列颠哥伦比亚大学; 范德堡大学; 宾夕法尼亚大学; 约翰霍普金斯大学; 纽约医学院; 不列颠哥伦比亚癌症中心; 德克萨斯大学 MD 安德森癌症中心; 康奈尔科技学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有病理学VLM基准的语言评估缺陷,提出以语言为中心的PathLang基准,通过固定图像相关要素仅改变诊断语言,评估发现VLMs对临床等价释义敏感,性能受语义竞争影响大。

AI 中文摘要

病理学视觉-语言模型(VLMs)已展现出强大的视觉感知能力,但其在语言领域的鲁棒性仍未得到充分表征。现有的病理学VLM基准大多依赖规范的闭集提示,或仅对通用模板进行扰动,将语言视为固定的评估组件,而非模型行为的可变轴。然而在临床实践中,诊断语言会因报告、机构及候选诊断的不同而存在差异。我们提出PathLang,这是一个以语言为中心且基于临床实际的零样本基准。PathLang固定了底层切片、真实标签及图像-文本评估方向,仅系统地改变诊断语言,从而使性能差异反映的是诊断的表述方式,而非图像内容的差异。语言变化遵循病理学家实际重新表述诊断的方式(术语、特异性及报告风格),所有提示和候选池均由6名持证病理学家验证。PathLang涵盖四类任务:(1)带图像-文本对齐分析的零样本分类;(2)跨模态检索;(3)释义鲁棒性,包括语义等价释义、长度与报告风格变化及提示集成;(4)在四个具有不同语义竞争形式的候选池上进行开放词汇诊断检索。在覆盖四个器官的五个公共数据集上的九个VLMs中,我们发现性能对临床等价释义高度敏感,在不同语义竞争形式间差异显著,且图像-文本对齐质量并不一定转化为类间可分性。我们在该httpsURL发布了提示语料库、候选池、预计算文本嵌入及评估代码。

英文摘要

Pathology vision-language models (VLMs) have shown strong visual perception ability, but their robustness in the language domain remains poorly characterized. Existing pathology VLM benchmarks largely rely on canonical closed-set prompts or perturb only generic templates, treating language as a fixed evaluation component rather than a variable axis of model behavior. In clinical practice, however, diagnostic language varies across reports, institutions, and candidate diagnoses. We introduce PathLang, a language-centered and clinically grounded zero-shot benchmark. PathLang holds the underlying slides, ground-truth labels, and image-text evaluation direction fixed while systematically varying only the diagnostic language, so that performance differences reflect how a diagnosis is phrased rather than what is imaged. The language variation follows how pathologists actually rephrase diagnoses (terminology, specificity, and reporting style), and all prompts and candidate pools are validated by six board-certified pathologists. PathLang covers four task families: (1) zero-shot classification with image-text alignment analysis, (2) cross-modal retrieval, (3) paraphrase robustness, including semantic-equivalence paraphrases, length and reporting-style variation, and prompt ensembling, and (4) open-vocabulary diagnosis retrieval over four candidate pools with distinct forms of semantic competition. Across nine VLMs and five public datasets spanning four organs, we find that performance is highly sensitive to clinically equivalent paraphrases, varies substantially across forms of semantic competition, and that image-text alignment quality does not necessarily translate into inter-class separability. We release the prompt corpus, candidate pools, pre-computed text embeddings, and evaluation code at https://anonymous.4open.science/r/PathLang.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑