AI 中文总结
本研究提出预注册单次最终测试评估框架,在三个转录组基础模型(Geneformer、scGPT、UCE)的外部数据集验证中,发现其均未可靠优于强表达基线,且属于一致的部分重复类别。
AI 中文摘要
转录组基础模型正日益被用作可复用的细胞与基因表征,但在弱监督和分布偏移的新数据上验证这类模型颇具挑战:标准对比方法会混淆真实表征信号与模型容量、行身份人工产物、对强任务特定基线的增益,以及在看到测试集后选定的结果规则。我们提出一种预注册的单次最终测试评估框架,该框架在查看任何测试数据前就锁定结果规则、随机种子和按目标基因分组的划分,并针对每个冻结表征评分,评分对象包括强表达基线、容量匹配的高斯对照,以及划分内的行身份(打乱)对照;仅每个细胞的嵌入提取步骤是模型特定的。将该框架应用于三个架构不同的模型——Geneformer、scGPT和UCE,涉及两个外部Replogle Perturb-seq数据集(RPE1和K562),结果显示三个模型均大幅通过了容量和行身份对照,但均未可靠地优于表达基线:最强的Geneformer最多仅超出基线约0.03的测试R²,且在两个数据集中均未达到预注册的五分之四种子阈值,而scGPT和UCE则低于基线。因此三个模型均落入相同的预注册部分重复类别——这是一种跨架构的一致结果,尽管各模型相对于基线的差距在符号和大小上存在差异。这些表征携带超出 trivial 对照的真实结构,但在这种弱量级标签下,无法迁移超越简单的强基线;该锁定框架可通过仅替换提取步骤,复用用于任何冻结转录组表征。
英文摘要
Transcriptomic foundation models are increasingly used as reusable cell and gene representations, but validating them on new data under weak supervision and distribution shift is hard: standard comparisons conflate genuine representation signal with model capacity, row-identity artifacts, gains over strong task-specific baselines, and outcome rules chosen after seeing the test set. We introduce a pre-registered, final-test-once evaluation framework that locks the outcome rule, seeds, and target-gene-grouped splits before any test data are seen, and scores each frozen representation against a strong expression baseline, a matched-capacity Gaussian control, and a within-split row-identity (shuffle) control; only the per-cell embedding-extraction step is model-specific. Applying it to three architecturally distinct models-Geneformer, scGPT, and UCE-across two external Replogle Perturb-seq datasets (RPE1 and K562), all three clear the capacity and row-identity controls by a wide margin, yet none reliably beats the expression baseline: the strongest (Geneformer) exceeds it by at most about $0.03$ test $R^2$ and clears the pre-registered four-of-five-seed threshold in neither dataset, while scGPT and UCE fall below it. All three therefore land in the same pre-registered partial-replication category-a consistent cross-architecture outcome, even though the baseline-relative gap differs in sign and magnitude across models. These representations carry real structure beyond trivial controls but, under this weak magnitude label, do not transfer past a simple strong baseline; the locked framework is reusable for any frozen transcriptomic representation by swapping only the extraction step.