arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

响应幅度作为保留型CRISPRi干扰效应预测的主导信号

Can We Trust In-Distribution Success? Locked Evaluation Reveals Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation Prediction

Mehrdad Shoeibi, Niloofar Yousefi

arXiv 2608.00152首次发表:更新:

发表机构

University of Central Florida(中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对保留型CRISPRi干扰效应预测,揭示响应幅度是主导信号,简单线性回归等模型性能优于深度编码器,且幅度特征可提升跨细胞类型迁移效果。

AI 中文摘要

预测CRISPRi干扰对保留型靶基因的转录组效应幅度是单细胞生物学中的重要开放问题。近期研究表明,在相关实验方案中,简单基线模型的表现往往与深度干扰预测模型相当甚至更优。我们在虚拟细胞挑战赛(Virtual Cell Challenge, VCC)基准数据集的严格保留型靶基因划分设置下研究该现象,识别出导致性能差距的特定低维信号,并表征其跨细胞类型的迁移特性。研究目标是靶基因与非靶向对照的对数Anderson-Darling距离,该距离可通过2000维输入的四个确定性标量函数进行强预测。直接访问完整输入的深度多层感知机(MLP)编码器会向边际训练均值坍缩,而标准补救措施无法缩小该差距。仅基于四个幅度标量的线性回归模型,其性能已超过最强的仅输入特征经典模型;将输入与四个幅度标量结合的随机森林模型,则显著优于我们的深度概念验证编码器。两个预先设定的对照实验将性能提升归因于逐行对齐而非增加维度。在针对两个外部CRISPRi筛选的零样本迁移任务中,采用从单细胞数据重建的靶基因端点进行评估,仅幅度预测器呈现正向迁移,而仅表达预测器则表现为负向或未确定。将幅度特征输入深度编码器可提升其相较于仅表达对应模型的迁移性能,但该编码器仍未优于基于相同特征的四标量线性回归模型。此外,我们发现随这些筛选提供的Anderson-Darling列指标衡量的是转录组范围的响应广度而非靶基因效应强度,因此针对该指标评估迁移会得到不同结果。

英文摘要

AI evaluation can support the wrong inference when an in-domain benchmark success does not survive distribution shift, or when the benchmark endpoint is entangled with a design factor. We study this problem in CRISPRi perturbation-effect prediction, evaluating a frozen Geneformer representation under a locked, pre-registered protocol: heads and model selection were frozen before test evaluation; the protocol required external outcome labels to remain withheld until final unblinding; and analysis-governing decisions were fixed before the evaluations they govern. In-distribution on the Virtual Cell Challenge (VCC), the frozen representation carries measurable predictive information beyond a dimension-matched random-feature control (Delta R^2 = +0.1645, 95% CI [+0.1375, +0.1920]), satisfying the pre-registered informativeness gate required before interpreting transfer. It then fails zero-shot transfer on both external screens (Spearman rho = -0.139 and -0.267), lying below that control on each. Adding a predefined magnitude block improves the representation externally (Delta rho = +0.032 and +0.143) but, under the frozen primary head, does not rescue transfer: both remain negative. A pre-registered, count-adjusted max-response secondary is positively associated with the outcome on both screens; we report it as correlational and secondary, not as a recovered magnitude signal. Finally, the VCC endpoint is strongly sample-size associated: a count-only linear model reaches R^2 = +0.4325, versus +0.2589 for the four magnitude scalars; adding those scalars to cell count improves R^2 by only +0.0017, so much of the aggregate-magnitude signal overlaps with cell count. This case study shows how locking the evaluation, harmonizing the measured endpoint, and separating primary from secondary evidence can change the inference supported by an AI benchmark.

Comments38 pages. Substantially revised version with updated framing, additional pre-specified controls and sensitivity analyses, expanded methodological and reproducibility documentation, and revised discussion of limitations. Main conclusions remain unchanged

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑