arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21060cs.AIcs.CV

CellPath-Bench:面向病理学基础模型的全片细胞表示多维基准

CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

Bokai Zhao, Yiyang Zhang, Hanqing Chao, Yawei Ma, Long Bai, Tai Ma, Minfeng Xu, Ming Song, Tianzi Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

CellPath-Bench是评估冻结病理学基础模型全片细胞表示能力的多维基准,通过对30款模型的测试揭示了细胞类型可解码性的模型差异,为相关审计提供标准化框架。

中文摘要 AI 辅助

病理学基础模型(PFMs)日益被用作通用骨干,但现有基准无法系统评估其全片细胞表示能力,包括细胞类型信息的可解码性,以及此类信息在组织切片、数据集和解剖器官间的可迁移性。我们推出CellPath-Bench,这是一款评估冻结PFMs的细胞分辨率基准。在对52个候选Xenium数据集进行质量控制后,我们构建了涵盖11个器官、包含7079283个细胞的25个空间对齐的H&E-Xenium组织切片面板,这些细胞被协调为细粒度和粗粒度分类体系。CellPath-Bench在配准的核坐标处采样冻结的WSI特征图,并使用标准化多类线性探针对其进行评估。细胞表示优势(CRA)衡量核锚定表示在切片内相对于patch级平均池化的优势,而细胞表示可迁移性(CRT)表征细胞类型可解码性在组织切片、数据集和器官间的泛化能力。我们通过空间读数、放大倍数、分类粒度和评估协议的304920次运行,对30个病理学专用和通用基础模型进行了基准测试。结果显示,细胞类型可解码性及其跨域泛化存在显著的模型依赖差异,产生了不同的多维能力轮廓。CellPath-Bench为审计冻结PFM表示中的细胞信息提供了标准化框架。

英文摘要

Pathology foundation models (PFMs) are increasingly used as general-purpose backbones, yet existing benchmarks cannot systematically diagnose their whole-slide cellular representation capabilities, including the decodability of cell-type information and the transferability of such information across tissue sections, datasets, and anatomical organs. We introduce CellPath-Bench, a cellular-resolution benchmark that evaluates frozen PFMs themselves. Following quality control of 52 candidate Xenium datasets, we construct a panel of 25 spatially aligned H\&E--Xenium tissue sections spanning 11 organs and 7,079,283 cells, harmonized into fine- and coarse-grained taxonomies. CellPath-Bench samples frozen WSI feature maps at registered nuclear coordinates and evaluates them using standardized multiclass linear probes. Cell Representation Advantage (CRA) measures the within-section advantage of nucleus-anchored representations over patch-level mean pooling, while Cell Representation Transferability (CRT) characterizes the generalization of cell-type decodability across tissue sections, datasets, and organs. We benchmark 30 pathology-specific and general-purpose foundation models through 304,920 runs across spatial readouts, magnifications, taxonomic granularities, and evaluation protocols. The results reveal substantial model-dependent differences in cell-type decodability and its cross-domain generalization, yielding distinct multidimensional capability profiles. CellPath-Bench provides a standardized framework for auditing cellular information in frozen PFM representations.

↑