arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07632q-bio.QMcs.CVeess.IV

JUMP-lite:紧凑、可复现的细胞表征基准测试

JUMP-lite: Compact, reproducible benchmarking of cell representations

  • Broad Institute of MIT and Harvard(麻省理工学院与哈佛大学博德研究所)

机构由 AI 辅助整理,请以论文原文为准。

Alán F. Muñoz, Johan Fredin Haslum, Runxi Shen, Anne E. Carpenter, Shantanu Singh

AI总结:

研究针对JUMP数据集体积过大导致细胞表征基准测试难以开展的问题,提出JUMP-lite压缩子集与Nahual框架,测试5种表征方法并验证压缩保留下游信号,为相关基准测试提供基础。

AI中文摘要:

基于图像的分析为药物发现和功能基因组学捕获丰富的表型特征。JUMP Cell Painting等大型公共数据集现提供数百万张图像用于系统研究,但JUMP数据集规模达115 TB,且评估实践分散,导致许多研究人员难以对表征方法进行系统比较。本文提出开源可复现模型部署框架Nahual,以及JUMP-lite——经整理的JUMP数据集116 GB子集,体积缩小1000倍,通过精心选择高置信度注释的扰动、采用有损JPEG XL压缩保留表型多样性。基于此,我们对5种表征方法进行基准测试,包括经典特征CellProfiler及深度学习模型MorphEM、OpenPhenom、SubCell、DINOv2,结果显示压缩保留了下游信号,标准化表型活性与一致性指标揭示了各方法间有意义的性能差异。JUMP-lite与Nahual为基于图像的细胞表征的可访问、可复现基准测试提供了基础。

英文摘要:

Image-based profiling captures rich phenotypic signatures for drug discovery and functional genomics. Large public datasets like JUMP Cell Painting now provide millions of images for systematic study. However, JUMP alone occupies 115 TB, and fragmented evaluation practices make systematic comparisons of representation methods impractical for many researchers. Here we present Nahual, an open-source framework for reproducible model deployment, and JUMP-lite, a 92.0 GB subset of JUMP that is approximately 1,250-fold smaller, selected to cover genetic modalities and compound annotations and reduced via lossy JPEG XL compression. Using these resources, we benchmark five representation methods, including classical features (CellProfiler) and deep learning models (MorphEM, OpenPhenom, SubCell, DINOv2). Moderate compression broadly retains signal relative to uncompressed images. Standardized phenotypic activity and consistency metrics reveal meaningful performance differences across methods. Together, JUMP-lite and Nahual provide a foundation for accessible, reproducible benchmarking of image-based cell representations.

补充信息

↑