发表机构
Stanford University; Genentech(斯坦福大学; 基因泰克)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出LLM引导的检索方法,在Tahoe-100M数据集上优于多个基线,可提升分子扰动预测的方向准确率,是零样本分子扰动预测的关键驱动因素。
AI 中文摘要
预测不同细胞系对小分子扰动的转录组响应是药物发现的核心,但穷尽式地分析药物-细胞组合不可行。我们将分子扰动预测构建为检索与聚合任务:通过聚合一小部分生物相关化合物的已测响应,近似某细胞系中未测药物的响应。我们提出大语言模型引导的检索(LLM-Guided Retrieval,LGR),其中大语言模型(Large Language Model,LLM)对候选邻域药物(限定为目标细胞系中已分析过的药物)进行排序,之后用固定均值聚合器组合这些药物的观测表达变化量以形成预测。我们在Tahoe-100M单细胞扰动图谱上,针对未见过的药物、未见过的细胞系和开放世界场景进行评估。LGR始终优于药物均值、ChemCPA和基于化学的kNN基线,在未见过的细胞系泛化中增益最强,其相关性高于均值基线、误差更低。在所有场景中,LGR提升了基因调控的方向(符号)准确率,表明即使基于幅度的指标相近,它也能更好地恢复具有生物学意义的扰动效应。这些结果显示,检索质量而非预测器复杂度是零样本分子扰动预测的关键驱动因素,且LLM作为受限检索模块时可提供有用的生物学先验。
英文摘要
Predicting transcriptomic responses to small-molecule perturbations across cell lines is central to drug discovery, but exhaustive profiling of drug-cell combinations is infeasible. We frame molecular perturbation prediction as retrieve-and-aggregate: approximate an unmeasured drug's response in a cell line by aggregating measured responses of a small set of biologically related compounds. We propose LLM-Guided Retrieval (LGR), where a large language model (LLM) ranks candidate neighbor drugs (restricted to those profiled in the target cell line); after which a fixed mean aggregator combines their observed expression deltas to form the prediction. We evaluate on the Tahoe-100M single-cell perturbation atlas under unseen-drug, unseen-cell-line, and open-world regimes. LGR consistently improves over drug mean, ChemCPA, and chemistry-based kNN baselines, with the strongest gains for unseen cell-line generalization, where it achieves higher correlation and lower error than mean baselines. Across settings, LGR improves directional (sign) accuracy of gene regulation, indicating better recovery of biologically meaningful perturbation effects even when magnitude-based metrics are similar. These results suggest that retrieval quality, rather than predictor complexity, is a key driver of zero-shot molecular perturbation prediction, and that LLMs can provide a useful biological prior when used as constrained retrieval modules.
CommentsPublished at MLGenX @ ICLR 2026