arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18578cs.CV

学习聚焦何处:用于组织病理学的自监督多尺度ViT

Learning Where to Focus: Self-Supervised Multi-Scale ViTs for Histopathology

Anabel Stammer, Valay Bundele, Mehran Hosseinzadeh, Hendrik P. A. Lensch

首次发表
浏览论文内容

中文总结 AI 辅助

提出CRAFT框架,通过自监督注意力学习在组织病理学图像中自适应分配空间分辨率,生成多尺度表示,以较低计算成本在多个基准上超越同等规模方法并媲美大型基础模型。

中文摘要 AI 辅助

病理学家通过首先定位可疑组织,然后在高倍率下检查来诊断疾病,而自监督视觉变换器(ViTs)则对每个图像区域分配相同的空间分辨率,尽管诊断证据稀疏且跨越多个生物尺度。最近的组织病理学基础模型通过扩大训练数据和模型容量显著提高了表示质量,但大多保留了统一的标记化。相反,我们研究是否可以通过在自监督学习期间学习分配空间分辨率的位置来改进组织病理学表示。为此,我们提出了CRAFT(从粗到细的区域自适应特征标记化),这是一种基于DINO的框架,通过使用自监督注意力选择性地细化信息区域同时保留粗略上下文,学习图像依赖的混合尺度表示,并配以对称的跨尺度正则化目标,以鼓励互补的粗略和精细表示。在CAMELYON16、TCGA-Lung亚型分类和TCGA-LUAD生存预测中,CRAFT在需要较低推理计算的情况下,始终优于同等规模的自监督方法。尽管仅使用紧凑的22M参数骨干网络,并在相对较小的组织病理学数据集上训练,CRAFT仍能与规模大得多的组织病理学基础模型保持竞争力,且常常超越它们。

英文摘要

Pathologists diagnose diseases by first locating suspicious tissue and then examining it at higher magnification, whereas self-supervised vision transformers (ViTs) allocate the same spatial resolution to every image region despite diagnostic evidence being sparse and spanning multiple biological scales. Recent pathology foundation models have substantially improved representation quality by scaling training data and model capacity, but largely retain uniform tokenization. We instead investigate whether pathology representations can be improved by learning where to allocate spatial resolution during self-supervised learning. To this end, we propose CRAFT (Coarse-to-fine Region-Adaptive Feature Tokenization), a DINO-based framework that learns image-dependent mixed-scale representations by using self-supervised attention to selectively refine informative regions while preserving coarse context, together with a symmetric cross-scale regularization objective that encourages complementary coarse and fine representations. Across CAMELYON16, TCGA-Lung subtype classification, and TCGA-LUAD survival prediction, CRAFT consistently outperforms comparable-scale self-supervised methods while requiring lower inference computation. Despite using only a compact 22M parameter backbone trained on comparatively small pathology datasets, CRAFT remains competitive with, and often surpasses, substantially larger pathology foundation models.

发表机构

  • Eberhard Karls Universität Tübingen(图宾根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑