arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32527cs.AIcs.LGq-bio.GNstat.AP

AmbiModBench:超越共享响应的基因扰动预测基准测试

AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses

Sikai Huang, Zhiwen Yang, Kai Yu, Jiayuan Chen, Stan Z. Li

AI总结:

针对基因扰动预测评估缺陷,提出AmbiModBench基准,通过配对参考、基因排列和嵌入距离分析,揭示绝对分数多反映共享背景,但RPE1上可检测到可重复的靶标特异性增益。

AI中文摘要:

预测细胞对基因扰动的反应有助于在单细胞基因组学中优先安排实验,因为在这些实验中,全面的测量是不可行的。尽管计算模型越来越多地用于预测这些反应,但三个评估缺陷掩盖了其分数所展示的内容。首先,绝对指标无法将靶标特异性预测与共享背景响应区分开来。其次,常见的指标在基因洗牌下仍然很高,因此基因水平的准确性从未得到验证。第三,在某一训练规模下的分数无法说明覆盖率,覆盖率取决于表示空间的邻近性和响应约束能力。我们提出了AmbiModBench,一个特异性感知、基因分辨率和覆盖率感知的基准。它将每个分数与在同一划分上拟合的训练均值参考配对,通过基因坐标排列筛选每个读数,并将嵌入距离与响应变异联系起来。在K562、RPE1和Norman数据集上,强绝对分数在很大程度上反映了共享背景而非靶标特异性学习。广泛使用的读数追踪的是响应幅度分布而非受影响的基因。可检测到的增益遵循表示空间的覆盖率而非训练集大小。尽管如此,在RPE1上,该协议在五个额外划分和三个基因选择中产生了可重复的靶标特异性增益,而仅凭绝对分数无法将其与共享背景区分开来。

英文摘要:

Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what their scores demonstrate. First, absolute metrics cannot separate target-specific predictions from a shared background response. Second, common metrics remain high under gene shuffling, so gene-level accuracy is never verified. Third, a score at one training size says nothing about coverage, which depends on representation-space proximity and response-constraining power. We propose AmbiModBench, a specificity-aware, gene-resolved and coverage-aware benchmark. It pairs every score with a training-mean reference fitted on the same split, screens each readout by gene-coordinate permutation, and links embedding distance to response variation. Across K562, RPE1 and Norman, strong absolute scores largely reflect shared background rather than target-specific learning. Widely used readouts track response magnitude distributions rather than the affected genes. Detectable gain follows representation-space coverage rather than training-set size. Nonetheless, on RPE1 the protocol yields a reproducible target-specific gain across five additional splits and three gene selections, which absolute scores alone cannot distinguish from shared background.

补充信息

↑