基因表达-FP:从观察到的细胞水平扰动特征中检索靶点和化合物
GeneSpeak-FP: Target and Compound Retrieval from Observed Cell-Level Perturbation Signatures
浏览论文内容
中文总结 AI 辅助
研究从观察到的细胞水平扰动特征中检索靶点和化合物的问题,提出Transformer检索模型\model,通过联合训练编码器映射特征,在Tahoe-100M条件下评估,结果表明该模型能从细胞反应中恢复靶点注释和化合物身份,但对未见情况的泛化待确定。
中文摘要 AI 辅助
大规模单细胞扰动图谱使得我们能够提出一个反向问题:给定一个观察到的转录反应,固定文库中的哪些注释靶点和化合物与该反应最一致?我们提出了一种名为\model的Transformer检索模型,用于这种封闭文库设置。每个输入是通过将一个处理过的细胞与细胞系特异性的平均DMSO对照进行对比而形成的细胞水平扰动特征。编码器将该特征映射到一个靶点检索向量和一个分子嵌入向量,通过监督靶点损失和结构-转录组比对进行联合训练。我们在Tahoe-100M条件下进行评估,使用10505个训练和1168个验证药物-细胞系对的化合物内分层90/10条件对分割。由于化合物和细胞系可能出现在两个分区中,实验测量的是留出条件对检索,而不是对未见化合物或细胞背景的泛化。在对38400个采样验证细胞的蒙特卡罗评估中,\model在379种化合物库上实现了靶点召回率@10为0.408,召回率@20为0.544,化合物命中率@1为0.129,命中率@10为0.343,平均倒数排名为0.205。单独的诊断评估显示主模型的值几乎相同,并且比随机向量控制和事后基因袋控制有很大提高。这些结果表明,在评估的Tahoe-100M封闭文库设置中,单个多任务模型可以从观察到的细胞水平反应中恢复映射的靶点注释和记录的化合物身份。对未见化合物和细胞背景的泛化仍有待确定。
英文摘要
Large-scale single-cell perturbation atlases make it possible to ask an inverse question: given an observed transcriptional response, which annotated targets and compounds in a fixed library are most consistent with that response? We present \model, a Transformer retrieval model for this closed-library setting. Each input is a cell-level perturbation signature formed by contrasting one treated cell with a cell-line-specific mean DMSO reference. The encoder maps the signature to a target-retrieval vector and a molecular-embedding vector, trained jointly with supervised target losses and structure--transcriptome alignment. We evaluate on Tahoe-100M conditions with mapped target annotations using a within-compound stratified 90/10 condition-pair split of 10,505 training and 1,168 validation drug--cell-line pairs. Because compounds and cell lines can occur in both partitions, the experiment measures held-out condition-pair retrieval rather than generalization to unseen compounds or cellular contexts. In a Monte Carlo evaluation over 38,400 sampled validation cells, \model\ achieved target Recall@10 of 0.408 and Recall@20 of 0.544, together with compound Hit@1 of 0.129, Hit@10 of 0.343, and mean reciprocal rank of 0.205 over a 379-compound bank. A separate diagnostic evaluation produced nearly identical values for the main model and large gains over a random-vector control and post-hoc bag-of-genes controls. These results demonstrate that a single multi-task model can recover both mapped target annotations and recorded compound identities from observed cell-level responses in the evaluated Tahoe-100M closed-library setting. Generalization to unseen compounds and cellular contexts remains to be established.
发表机构
- AIFFEL Research, Modulabs(艾菲研究公司,模块实验室)
机构由 AI 辅助整理,请以论文原文为准。