单细胞基础模型基准中的预训练污染审计
Auditing pretraining contamination in single-cell foundation model benchmarks
AI总结:
研究单细胞基础模型基准中的预训练污染问题,提出scContam审计框架,结合指纹信号与成员推理攻击。应用于多个基准和模型,发现部分基准有污染,且MIA-scFM的AUROC与模型容量数据比有关,证明预训练审计可行且应伴随基准报告。
AI中文摘要:
诸如Geneformer、scGPT和通用细胞嵌入(UCE)等单细胞基础模型(scFMs)在从公共存储库提取的数千万个细胞上进行预训练。相同的存储库也是广泛使用的整合基准的基础,这带来了一种未被衡量的风险,即零样本基准性能反映的是预训练暴露而非真正的泛化能力。我们引入了scContam,这是一个针对每个细胞的审计框架,它将基于MinHash的基因集指纹信号与明确的预训练语料库相结合,并采用基于损失的成员推理攻击(MIA-scFM)。应用于四个scIB基准和三个scFMs,我们发现两个引用最多的基准,PBMC 3k和CELLxGENE人类胰岛图谱,包含大量预训练重叠证据(分别有80.4%和77.0%的细胞与Genecorpus-30M的指纹p < 0.05),而后截止数据集AIDA v2和Tahoe-100M没有重叠证据(0%)。一个受控的重新预训练实验表明,MIA-scFM的曲线下面积(AUROC)随着模型的容量与数据比率单调增加(在适当正则化、轻度过拟合和激进过拟合状态下,AUROC分别为0.494→0.690→0.881),这表明生产中的scFMs能够抵抗实例级记忆,但分布污染必须单独检测。一项使用三种架构的供体匹配的细胞类型内分析表明,受污染的细胞比供体匹配的干净细胞嵌入得更紧密(排列p值分别为0.030、0.014、< 0.002),AIDA阴性对照完全为零。预训练审计是可行的,应伴随scFM基准报告。
英文摘要:
Single-cell foundation models (scFMs) such as Geneformer, scGPT, and Universal Cell Embeddings (UCE) are pretrained on tens of millions of cells drawn from public repositories. The same repositories underlie widely used integration benchmarks, creating an unmeasured risk that zero-shot benchmark performance reflects pretraining exposure rather than genuine generalization. We introduce \textbf{scContam}, a per-cell audit framework that combines a MinHash-based gene-set fingerprint signal against the explicit pretraining corpus with a loss-based membership inference attack (MIA-scFM). Applied to four scIB benchmarks and three scFMs, we find that two of the most-cited benchmarks, PBMC 3k and the CELLxGENE human pancreatic islet atlas, contain extensive pretraining-overlap evidence ($80.4\%$ and $77.0\%$ of cells with fingerprint $p < 0.05$ against Genecorpus-30M), whereas the post-cutoff datasets AIDA v2 and Tahoe-100M show no overlap evidence ($0\%$). A controlled re-pretraining experiment establishes that MIA-scFM AUROC scales monotonically with the model's capacity-to-data ratio (AUROC $0.494 \to 0.690 \to 0.881$ across properly-regularized, mildly-overfit, and aggressively-overfit regimes), demonstrating that production scFMs resist instance-level memorization but distributional contamination must be detected separately. A donor-matched, within-cell-type analysis with three architectures shows that contaminated cells embed measurably more tightly than donor-matched clean cells (permutation $p = 0.030, 0.014, < 0.002$, respectively), with a perfectly null AIDA negative control. Pretraining audits are tractable and should accompany scFM benchmark reporting.