AI 中文总结
本文针对病理VLM的证据使用问题,提出CleanSlide基准和Pair-DPO方法,通过反事实配对训练提升图像证据利用,显著优于现有模型。
AI 中文摘要
病理学视觉语言模型(VLM)通常通过准确率进行评估,但准确率本身并不能衡量证据的使用:它可能混淆数据集污染、先验知识和图像证据。在一项关于淋巴结转移预测的动机研究中,我们发现大多数公开的病理学VLM在输入全切片图像、标注病灶或不输入图像时,表现差异极小。为了更好地理解这些模型所利用的具体特征,本文提出了两项贡献,旨在分解这些因素。首先,我们提出了CleanSlide,一个基于TCGA的VQA基准,旨在消除图像侧和问题侧的污染。它包含149K个经过审计的多项选择题,覆盖9,985张切片,并采用患者和组织来源不重叠的划分。每个问题都经过选项捷径、题干泄漏、跨划分重复和盲解性的审计。其次,我们提出了Pair-DPO,一种基于同一问题和来源的反事实切片对的偏好损失。通过控制共享的混淆因素,Pair-DPO消除了问题归因信号,使图像证据成为偏好的来源。具体来说,每一对包含两张具有相反、已验证发现的真实切片,不引入编辑伪影或未验证的标签,适用于弥漫性或分级特征,如浸润、坏死和肿瘤分级。实验表明,我们的方法在CleanSlide上从图像证据中获得了15.29%的提升,而最佳已发表模型仅为2.81%。在外部CPTAC和BCNB队列上,我们的方法分别达到了57.6%和59.0%的准确率,超过所有其他评估模型9.7%和3.4%。我们将发布基准和代码。
英文摘要
Pathology vision-language models (VLMs) are conventionally evaluated by accuracy, but accuracy alone does not measure evidence use: it may conflate dataset contamination, prior knowledge, and image evidence. In a motivating study of lymph-node metastasis prediction, we found that most public pathology VLMs showed minimal differences when changing from feeding the models with whole-slide images, an annotated lesion, or no image at all. To better understand the specific features leveraged by these models, this paper presents two contributions aimed at disentangling these factors. First, we present CleanSlide, a TCGA-based VQA benchmark designed to eliminate image- and question-side contamination. It contains 149K audited multiple-choice questions over 9,985 slides, with patient- and tissue-source-disjoint splits. Every question is audited for option shortcuts, stem leakage, cross-split duplication, and blind solvability. Second, we propose Pair-DPO, a preference loss over counterfactual slide pairs from the same question and source. By controlling for shared confounding factors, Pair-DPO cancels out the question-attributable signal and leaves image evidence as the source of preference. Specifically, each pair consists of two real slides with opposite, verified findings, introducing neither editing artifacts nor unverified labels for diffuse or graded features such as invasion, necrosis, and tumor grade. Experiments show that our method gains 15.29% from image evidence on the CleanSlide, compared with 2.81% for the best published model. On the external CPTAC and BCNB cohorts, our method achieves accuracies of 57.6% and 59.0%, outperforming all other evaluated models by 9.7% and 3.4%, respectively. We will release the benchmark and code.
Comments19 pages, 5 figures, 6 tables