arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12761cs.CVcs.LGq-bio.QM

相同编码器,不同胜者:用于细胞绘画编码器评估的配对视图框架

Same Encoder, Different Winner: A Paired-View Framework for Cell Painting Encoder Evaluation

  • Institute of Computational Biology, Helmholtz Munich(亥姆霍兹慕尼黑计算生物学研究所)
  • School of Computation, Information and Technology, Technical University of Munich(慕尼黑工业大学计算、信息与技术学院)
  • Broad Institute of MIT and Harvard(麻省理工学院和哈佛大学布罗德研究所)

机构由 AI 辅助整理,请以论文原文为准。

Tim Treis, Nikita Moshkov, Johan Fredin Haslum, Shantanu Singh, Fabian J. Theis

AI总结:

提出CP-BG-Bench配对视图框架,通过固定中心细胞并消融/增强背景,揭示细胞绘画编码器排名随评估协议系统变化,分歧可归因于细胞-背景、形态-上下文和研究内-跨批次三个轴,背景增益由实验设计决定。

AI中文摘要:

细胞绘画的视觉编码器通常通过单一评估进行排名,常见的是复现平均精度(mAP)。我们引入了CP-BG-Bench,一个配对视图评估框架,该框架在四个匹配视图(原始裁剪C、分割S以及密度增强变体CD和SD)中保持中心细胞固定,通过消融或增强周围像素作为受控干预。将该框架实例化于三个数据集(JUMP-CP、RxRx1、RxRx3-core)和三个编码器(DINOv3 ViT-B/16、OpenPhenom、SubCell)上,并采用四种社区标准协议(复现mAP、scIB批次整合、CellProfiler特征预测、跨批次扰动召回),我们发现四种协议对相同编码器的排名系统性地不同,分歧沿三个轴分解:细胞与背景、形态与上下文、研究内与跨批次。最大的效应:在RxRx3-core上,使用分割输入的SubCell保留了裁剪复现mAP的94%,但仅保留了裁剪R@10的32%,因此分割下保留的研究内信号在很大程度上不可转移;密度增强恢复了研究内C到S差距的84%,但仅恢复了跨批次差距的8%。分割视图在三个数据集中的两个上预测CellProfiler特征与裁剪相当或更好,逆转了复现mAP排名,且C到S差距在数据集间变化达一个数量级,而在编码器间保持相似,表明背景驱动的增益由实验设计而非编码器决定。因此,细胞绘画编码器的单一指标排名对所用协议敏感,协议分歧可解释为配对视图设计所暴露的三个轴上的投影。我们将发布配对视图数据集、重建流程、36个训练检查点、聚合嵌入以及完整评估套件。

英文摘要:

Vision encoders for Cell Painting are typically ranked by a single evaluation, commonly replicate mean average precision (mAP). We introduce CP-BG-Bench, a paired-view evaluation framework that holds the central cell fixed across four matched views (raw crop C, segmented S, and density-augmented variants CD and SD), ablating or augmenting surrounding pixels as a controlled intervention. Instantiating the framework on three datasets (JUMP-CP, RxRx1, RxRx3-core) and three encoders (DINOv3 ViT-B/16, OpenPhenom, SubCell) under four community-standard protocols (replicate mAP, scIB batch integration, CellProfiler feature prediction, cross-batch perturbation recall), we find that the four protocols rank the same encoders systematically differently, with disagreements decomposing along three axes: cell versus background, morphology versus context, and within-study versus across-batch. The largest effect: on RxRx3-core, SubCell with segmented inputs retains 94% of crop replicate mAP but only 32% of crop R@10, so the within-study signal preserved under segmentation is largely non-transferable; density augmentation recovers 84% of the within-study C-to-S gap but only 8% of the cross-batch gap. Segmented views predict CellProfiler features as well as or better than crops on two of three datasets, inverting the replicate-mAP ranking, and the C-to-S gap varies by an order of magnitude across datasets while remaining similar across encoders, indicating that background-driven gain is set by experimental design rather than by the encoder. Single-metric ranking of Cell Painting encoders is therefore sensitive to the protocol used, and protocol disagreements are interpretable as projections onto the three axes the paired-view design exposes. We will release the paired-view datasets, reconstruction pipelines, 36 trained checkpoints, aggregated embeddings, and the full evaluation suite.

补充信息

↑