发表机构
The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对病理视觉推理的证据基础评估缺口,构建PathoArgus-Bench基准与ESG受控集,发现现有模型准确率高但证据基础弱,提出PathoArgus阅读器,呼吁转向证据导向的评估。
AI 中文摘要
全切片病理推理要求模型整合千兆像素级的视觉证据与完整病例关联切片,但现有问答基准主要衡量最终答案准确率——该指标易受语言先验和基准规律影响,且不足以证明预测是基于提供的组织。我们推出PathoArgus-Bench,这是一个明确测试完整证据链(可用性、可访问性、使用及响应性)的基准与评估协议。PathoArgus-Bench包含来自TCGA的15个项目中4913名患者的22078道四选一问题,涵盖三个证据需求层级的六项病理能力,并在固定阅读器预算下运行,仅保留千兆像素语境的一小部分。为进一步隔离基于证据的推理,我们提出ESG(证据状态四重奏),这是一套受控的483个四重奏,其中问题文本固定,而目标WSI集被移动、替换或移除,要求在所有状态下保持一致预测。对20个通用、医疗及病理专用系统的评估显示出巨大差距:GPT-5.6的总体准确率达57.09%,在ESG上的准确率为57.04%,但仅正确完成483个四重奏中的19个(3.93%的QExact),暴露出行级准确率无法转化为可靠的证据基础。我们还推出PathoArgus,一种固定预算阅读器,通过问题相关性和空间覆盖分配语境,总体准确率达50.39%,但QExact仅为1.86%——证明仅改进语境访问不足以确保一致的基于证据的预测。我们的基准与诊断结果表明,获取有用的全切片语境是必要的,但远非充分,并呼吁计算病理学从以答案为中心转向以证据为中心的评估。
英文摘要
Whole-slide pathology reasoning requires models to integrate gigapixel-scale visual evidence across complete case-linked slides, yet current question-answering benchmarks primarily measure final answer accuracy--a metric vulnerable to linguistic priors and benchmark regularities, and insufficient to establish that predictions are grounded in the supplied tissue. We introduce PathoArgus-Bench, a benchmark and evaluation protocol that explicitly tests the full evidence chain: availability, accessibility, use, and responsiveness. PathoArgus-Bench comprises 22,078 four-choice questions from 4,913 patients across 15 TCGA projects, covering six pathology capabilities across three levels of evidence demand, and operates under a fixed reader budget that retains only a small fraction of the gigapixel context. To further isolate evidence-grounded reasoning, we contribute ESG (Evidence State Quartets), a controlled set of 483 quartets where the question text is fixed while the target WSI set is moved, replaced, or removed, requiring consistent predictions across all states. Evaluating 20 general-purpose, medical, and pathology-specific systems reveals a stark gap: while GPT-5.6 achieves 57.09% overall accuracy and 57.04% on ESG, it correctly completes only 19 of 483 quartets (3.93% QExact), exposing that row-level accuracy does not translate into reliable evidence grounding. We also introduce PathoArgus, a fixed-budget reader that allocates context via question relevance and spatial coverage, attaining 50.39% overall accuracy yet only 1.86% QExact--demonstrating that improved context access alone does not ensure consistent evidence-based prediction. Our benchmark and diagnostics establish that acquiring useful whole-slide context is necessary but far from sufficient, and call for a shift from answer-centric to evidence-grounded evaluation in computational pathology.