多重实例学习的统计检验:基于选择性推断及其在计算病理学中的应用
Statistical Testing for Multiple Instance Learning via Selective Inference with Applications to Computational Pathology
浏览论文内容
中文总结 AI 辅助
本文提出基于选择性推断的统计检验框架,用于评估注意力MIL中高注意力实例的显著性,在合成、MNIST及病理WSI数据上实现I型错误控制并提升统计功效。
中文摘要 AI 辅助
多重实例学习(MIL)在计算病理学中被广泛使用,因为它能够在无需补丁级标注的情况下对全切片图像(WSIs)进行弱监督分析。在基于注意力的MIL中,具有高注意力分数的实例通常被解释为诊断上重要的区域,并用作视觉解释。然而,仅凭注意力分数无法确定所选择的高注意力实例是否与正常实例显著不同,这限制了基于注意力的解释的可靠性。在本文中,我们将高注意力实例的评估表述为一个统计假设检验问题。具体来说,我们评估所选择的高注意力实例是否显著偏离基于特征相似性选择的代表性正常参考实例。一个主要挑战是,目标实例和参考实例都是通过数据依赖的过程选择的,这使得标准假设检验失效。为了解决这个问题,我们引入了一个选择性推断(SI)框架,该框架明确考虑了由基于注意力的实例选择和自适应参考选择引发的选择事件,从而能够在这些事件条件下计算有效的选择性$p$值。实验表明,在合成数据和基于MNIST的数据上控制了I型错误,并且在实际病理WSIs上具有实用性,其统计功效高于传统的过度条件化方法。
英文摘要
Multiple instance learning (MIL) is widely used in computational pathology because it enables weakly supervised analysis of whole-slide images (WSIs) without requiring patch-level annotations. In attention-based MIL, instances with high attention scores are often interpreted as diagnostically important regions and used as visual explanations. However, attention scores alone cannot determine whether selected high-attention instances are significantly different from normal instances, limiting the reliability of attention-based explanations. In this paper, we formulate the evaluation of high-attention instances as a statistical hypothesis testing problem. Specifically, we assess whether a selected high-attention instance significantly deviates from a representative normal reference instance selected based on feature similarity. A major challenge is that both the target instance and the reference instance are selected through data-dependent procedures, rendering standard hypothesis testing invalid. To address this issue, we introduce a selective inference (SI) framework that explicitly accounts for the selection events induced by attention-based instance selection and adaptive reference selection, thereby enabling the computation of valid selective $p$-values conditional on these events. Experiments demonstrate Type-I error control on synthetic and MNIST-based data and practical applicability to pathological WSIs, with higher statistical power than the conventional over-conditioning approach.