arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

全切片图像多模态基准测试中的数据泄露审计

Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks

Wenhao Zhang, Zhongliang Zhou, John Kang, Sheng Li

arXiv 2607.12278首次发表:更新:

发表机构

University of Virginia; Merck & Co., Inc.(弗吉尼亚大学; 默克公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对计算病理学中WSI VQA基准测试数据泄露问题,通过追踪标识符发现患者级和机构级泄露,证明其可从基础模型特征空间解码,导致准确率差距,当前评估无法区分推理与检索,最后给出无污染评估建议。

AI 中文摘要

近期用于计算病理学的视觉语言模型(VLM)在全切片图像(WSI)视觉问答(VQA)基准测试中展现出惊人的零样本性能。我们对这些结果进行审计,发现它们在两个层次上存在数据泄露问题:患者级泄露,即同一病例的切片出现在训练集和测试集中;机构级泄露,即不同病例通过共同的组织源站点(TSS)共享染色批次和扫描仪特征。通过追踪主要公共资源中的切片、病例和TSS标识符,我们记录了基于TCGA的基准测试中病例级训练测试重叠率为92.3%至100%,以及几乎完全的TSS重叠。我们进一步证明,这两种泄露水平都可以从基础模型特征空间中线性解码,它们在已发布的检查点上导致了泄露病例和审计清理病例之间可测量的准确性差距,并且在多个已发布的WSI VLM中,报告的峰值准确率集中在污染最严重的基准测试上。因此,当前的WSI VQA评估无法区分真正的多模态推理和基于记忆的机构及患者特定工件的最近邻检索。最后,我们概述了无污染评估的具体建议。通过解决基准构建、来源披露和自动重叠审计问题,我们旨在引导未来的研究朝着可验证的进展声明发展。

英文摘要

Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answering (VQA) benchmarks. We audit these claims and find them fundamentally compromised by data leakage at two hierarchical levels: patient-level leakage, where slides from the same case appear in both training and test folds, and institutional-level leakage, where different cases nonetheless share staining-batch and scanner signatures through a common Tissue Source Site (TSS). By tracing canonical slide, case, and TSS identifiers across major public resources, we document case level train test overlaps of 92.3~100% on TCGA-derived benchmarks, together with near-complete TSS overlap. We further demonstrate that both leakage levels are linearly decodable from foundation-model feature space, that they induce a measurable accuracy gap between leaked and audit-clean cases on a published checkpoint, and that across multiple published WSI VLMs, peak reported accuracies concentrate on the most heavily contaminated benchmarks. Therefore, the current WSI VQA evaluation cannot distinguish genuine multimodal reasoning from nearest-neighbor retrieval over memorized institutional and patient-specific artifacts. Finally, we outline concrete recommendations for contamination-free evaluation. By addressing benchmark construction, provenance disclosure, and automated overlap auditing, we aim to guide future research toward verifiable claims of progress.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑