arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一瓦片,多实例:重新思考稀疏诊断证据下的多实例学习

One Tile, Multiple Instances: Rethinking MIL for Sparse Diagnostic Evidence

Runsheng Liu, Cheng Jin, Hao Jiang, Hao Chen

arXiv 2610.04853首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute; State Key Laboratory of Nervous System Disorders; SIAT-HKUST Joint Laboratory for Brain Science(香港科技大学; 香港科技大学深圳-香港协同创新研究院; 神经系统疾病国家重点实验室; 深圳先进技术研究院-香港科技大学脑科学联合实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对弱监督WSI分类中瓦片级嵌入掩盖细粒度证据的问题,提出DI-MIL框架,通过聚类瓦片内空间令牌生成多实例嵌入,无需训练即可提升稀疏诊断场景下的分类性能。

AI 中文摘要

在弱监督全切片图像(WSI)分类中,特征提取器通常将每个图像瓦片压缩为单个全局嵌入。因此,切片级聚合器被限制在这种粗粒度的瓦片尺度上,从而向注意力机制隐藏了细粒度的子瓦片证据。我们提出了DI-MIL,一个通过分解实例将编码上下文与实例粒度解耦的框架。通过在每片瓦片内对来自冻结基础模型的密集空间令牌进行聚类,DI-MIL将单个瓦片转换为多个独立加权的实例嵌入。作为一个无需训练的编码后模块,DI-MIL可无缝集成到现有流程中,无需重新编码或修改下游架构。我们在细胞病理学上评估了DI-MIL,这是一个具有挑战性的测试平台,其中稀疏的诊断信号在标准瓦片中容易被稀释。在四个数据集、三个冻结基础模型和两个基于注意力的聚合器上,DI-MIL表现出一致的有效性,在72项指标级比较中改善了67项,在细胞病理学专用骨干下最大平均增益达到3.64个百分点。在与七个代表性MIL基线的更广泛比较中,DI-MIL与ACMIL配对在33项骨干-数据集-指标比较中取得了最高平均性能。消融研究表明,直接使用更小的瓦片会将提取的瓦片数量膨胀至多43.3倍,且性能非单调,而DI-MIL在实现最强总体结果的同时,零额外图像提取开销。这些结果确立了实例构建作为MIL中一个正交设计维度,支持DI-MIL作为稀疏诊断证据下的高性价比解决方案。

英文摘要

In weakly supervised Whole Slide Image (WSI) classification, feature extractors typically compress each image tile into a single global embedding. Consequently, slide-level aggregators are restricted to this coarse tile scale, concealing fine-grained sub-tile evidence from the attention mechanism. We introduce DI-MIL, a framework that decouples encoding context from instance granularity through decomposed instances. By clustering dense spatial tokens from a frozen foundation model within each tile, DI-MIL converts a single tile into multiple independently weighted instance embeddings. As a training-free post-encoding module, DI-MIL integrates seamlessly into existing pipelines without requiring re-encoding or downstream architectural modifications. We evaluate DI-MIL on cytopathology, a challenging testbed where sparse diagnostic signals are easily diluted within standard tiles. Across four datasets, three frozen foundation models, and two attention-based aggregators, DI-MIL demonstrates consistent efficacy, improving 67 of 72 metric-level comparisons, with the largest mean gains reaching 3.64 points under cytopathology-specific backbones. In a broader comparison against seven representative MIL baselines, DI-MIL paired with ACMIL achieves highest mean performance in 33 of 36 backbone-dataset-metric comparisons. Ablations show that direct smaller tiling inflates the extracted tile count by up to 43.3$\times$ with non-monotonic performance, whereas DI-MIL incurs zero additional image-extraction overhead while achieving the strongest overall results. These results establish instance construction as an orthogonal design dimension in MIL, supporting DI-MIL as a cost-efficient solution under sparse diagnostic evidence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑