AI 中文总结
针对现有计算病理学模型泛化性受限的问题,提出多分辨率金字塔Transformer(MRPT),经多分辨率自监督预训练后,在34个数据集的多项任务中性能优于同类模型与多模态大语言模型。
AI 中文摘要
视觉Transformer(ViTs)及其分层变体在计算病理学(CPath)中已取得优异性能,但多数模型在单分辨率全切片图像(WSIs)上进行预训练,限制了其在任意分辨率间的泛化能力。千兆像素WSIs天然包含多尺度诊断模式,涵盖细胞形态、组织结构与全局上下文,与病理专家检查WSIs的方式一致。本文提出多分辨率金字塔Transformer(MRPT),该模型可分层聚合从细胞到组织再到WSI级别的多分辨率信息,采用具有生物学意义的连续跨分辨率注意力(CCRA)机制捕捉尺度无关交互,通过对齐不同分辨率的嵌入来强化多分辨率语义一致性,从而生成鲁棒且泛化性强的WSI表征。MRPT在6.24亿个图像块、240万个区域及3.6万张WSIs上以多分辨率自监督方式预训练,学习到丰富的从粗到细的组织病理学特征。在34个多样化数据集上开展的大量实验表明,MRPT在癌症亚型分类、组织表型分型及WSI理解的视觉问答(VQA)任务中,性能优于近期的基础模型与多模态大语言模型(MLLMs)。
英文摘要
Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution Whole Slide Images (WSIs), limiting their generalization across arbitrary resolutions. Gigapixel WSIs inherently contain diagnostic patterns at multiple scales, including cellular morphologies, tissue architectures, and global context, mirroring how expert pathologists examine WSIs. We introduce Multi-Resolution Pyramid Transformer (MRPT), a model that hierarchically aggregates multi-resolution information from cellular to tissue and WSI levels. MRPT employs a biologically meaningful Consecutive Cross-Resolution Attention (CCRA) mechanism to capture scale-independent interactions and enforces multi-resolution semantic consistency by aligning embeddings across resolutions, yielding robust and generalizable WSI representations. Pre-trained in a multi-resolution self-supervised manner on 624M patches, 2.4M regions, and 36K WSIs, MRPT learns rich coarse-to-fine histopathology features. Extensive experiments on 34 diverse datasets show that MRPT surpasses recent foundation models and Multimodal Large Language Models (MLLMs) in cancer subtype classification, tissue phenotyping, and Visual Question Answering (VQA) for WSI understanding.