从图像块到证据球:面向小样本全切片图像分类的类别条件证据检索
From Patches to Evidence Balls: Class-Conditioned Evidence Retrieval for Few-Shot Whole Slide Image Classification
浏览论文内容
中文总结 AI 辅助
该研究针对小样本全切片图像分类任务,提出类别条件证据检索框架EviBall,通过组织证据球并结合任务特定类别查询,提升分类性能与可解释性,在四项相关任务上优于现有基线方法。
中文摘要 AI 辅助
全切片图像(Whole Slide Image, WSI)分类是一项证据驱动的任务,其诊断线索通常较为稀疏、具有空间组织性且与类别相关。现有的多实例学习(Multiple Instance Learning, MIL)及视觉-语言方法会将大量图像块特征聚合为单一的全局切片表示。在小样本监督场景下,有限的切片级标签使得难以学习可靠的聚合机制,该机制需将稀疏的局部线索组织为紧凑且连贯的诊断证据。此外,共享的切片表示会将支持候选类别及其替代类别的证据压缩至同一特征中,限制了类别特定的推理与可解释性。为解决这些问题,我们提出EviBall,这是一种面向小样本WSI分类的类别条件证据检索框架。EviBall通过语义-空间分配与中心优化将局部图像块组织为证据球,在弱监督下生成紧凑且空间连贯的证据单元。随后,它利用特定任务的类别查询(包括面向形态学任务的语言引导查询、面向分子终点预测的分子引导查询)来检索支持性证据球,并生成用于直接类别预测的类别条件证据表示。通过引入结构化证据单元与任务相关的语义引导,EviBall减少了对从稀缺切片级标签学习无约束全局聚合机制的依赖,从而将小样本WSI分类重新表述为候选类别间的结构化证据检索与竞争。在四项面向形态学与分子终点的WSI任务上开展的大量实验表明,EviBall在各类小样本设置下均始终优于传统及视觉-语言MIL基线,同时为每次预测提供空间定位且类别特定的证据。
英文摘要
Whole slide image (WSI) classification is an evidence-driven task, where diagnostic cues are often sparse, spatially organized, and class-dependent. Existing MIL and vision-language methods aggregate a large pool of patch features into a single global slide representation. Under few-shot supervision, limited slide-level labels make it difficult to learn a reliable aggregation mechanism that organizes sparse local cues into compact and coherent diagnostic evidence. Moreover, a shared slide representation compresses evidence supporting a candidate class and its alternatives into the same feature, limiting class-specific reasoning and interpretability. To address these issues, we propose EviBall, a class-conditioned evidence retrieval framework for few-shot WSI classification. EviBall organizes local patches into Evidence Balls through semantic-spatial assignment and center refinement, yielding compact and spatially coherent evidence units under weak supervision. It then uses task-specific class queries, including language-guided queries for morphology-oriented tasks and molecular-guided queries for molecular endpoint prediction, to retrieve supporting evidence balls and produce class-conditioned evidence representations for direct class-wise prediction. By introducing structured evidence units and task-relevant semantic guidance, EviBall reduces the reliance on learning an unconstrained global aggregation mechanism from scarce slide-level labels. It therefore reformulates few-shot WSI classification as structured evidence retrieval and competition among candidate classes. Extensive experiments across four morphology-oriented and molecular endpoint WSI tasks demonstrate that EviBall consistently outperforms conventional and vision-language MIL baselines under diverse few-shot settings, while providing spatially localized and class-specific evidence for each prediction.
发表机构
- XJTU(西安交通大学)
- JUFE(江西财经大学)
- University of Cambridge(剑桥大学)
- NUS(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。