arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26382cs.CV

VIPER:用于兽医病理学视觉语言模型的专家策划基准

VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology

Luca L. Weishaupt, Simone de Brot, Javier Asin, Llorenç Grau-Roma, Nic G. Reitsam, Andrew H. Song, Dongmin Bang, Stefan T. Kaluziak, Long Phi Le, Jakob Nikolas … 展开作者

Luca L. Weishaupt, Simone de Brot, Javier Asin, Llorenç Grau-Roma, Nic G. Reitsam, Andrew H. Song, Dongmin Bang, Stefan T. Kaluziak, Long Phi Le, Jakob Nikolas Kather, Faisal Mahmood, Guillaume Jaume

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对现有病理视觉语言模型基准聚焦人体组织的问题,推出首个兽医病理学视觉语言模型评估基准VIPER,测试16类模型发现领域差距,验证领域特定训练的重要性。

中文摘要 AI 辅助

病理学视觉语言模型发展迅速,但现有基准仍聚焦于人体组织,尤其是肿瘤学,未涉及非人病理学。这一空白在毒理学病理学中尤为重要,对实验动物进行显微镜组织检查是临床前药物安全性评估的核心部分。为解决该问题,我们推出VIPER,即首个用于毒理学病理学视觉语言模型评估的专家策划基准。VIPER包含1251个问题,关联419张HE染色大鼠组织学图像,涵盖7个器官系统,题型包括选择题、KPrim型和自由文本型。所有问题均经认证兽医病理学家策划与验证。我们共对16个模型进行基准测试,其中包括2个新推出的兽医病理学模型、7个人体病理学专用模型和7个通用前沿模型。结果显示,兽医与人体病理学间存在显著领域差距,前沿模型存在正常组织过度诊断的风险,且领域特定训练对视觉基础预测仍至关重要。VIPER数据与评估代码可在指定URL获取。

英文摘要

Pathology vision-language models are advancing rapidly, yet existing benchmarks remain focused on human tissue, particularly oncology, leaving non-human pathology largely unaddressed. This gap is especially important in toxicologic pathology, where microscopic tissue examination of laboratory animals is a core component of preclinical drug safety assessment. To address it, we introduce VIPER, the first expert-curated benchmark for vision-language model evaluation in toxicologic pathology. VIPER contains 1,251 questions associated with 419 H&E-stained rat histology images across seven organ systems, covering multiple-choice, KPrim, and free-text formats. All questions were curated and validated by board-certified veterinary pathologists. In total, we benchmarked 16 models, including two newly introduced veterinary-pathology models, seven human pathology-specialized models, and seven general-purpose frontier models. The results identify a substantial domain gap between veterinary and human pathology, expose the risk of over-diagnosis of normal tissue in frontier models, and show that domain-specific training remains critical for visually grounded predictions. VIPER data and evaluation code are available at https://github.com/mahmoodlab/viper.

发表机构

  • Harvard-MIT HST(哈佛-麻省理工卫生科学与技术部)
  • Mass General Brigham(麻省总医院布里格姆医疗系统)
  • Harvard Medical School(哈佛医学院)
  • COMPATH, University of Bern(伯尔尼大学COMPATH)
  • UC Davis(加州大学戴维斯分校)
  • University of Augsburg(奥格斯堡大学)
  • UT MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)
  • TU Dresden(德累斯顿工业大学)
  • University of Lausanne(洛桑大学)

机构由 AI 辅助整理,请以论文原文为准。

↑