发表机构
University of Cambridge; Prior Labs(剑桥大学; 普瑞尔实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对预测细胞对未知化学扰动反应的挑战,提出PerturbPFN模型,通过推断潜在系统图等并经SCM解码器传播效果,基于合成情节训练,在真实和合成数据上评估,实现低推理成本且具可解释性的扰动预测。
AI 中文摘要
预测细胞对未知化学扰动的反应具有挑战性,原因包括未知靶点和机制、高维表达反应以及小分子设计空间实验覆盖有限。我们提出PerturbPFN,一种在分层合成结构先验下用于未知靶点扰动预测的PFN风格摊销模型。它不直接回归高维表达反应,而是推断潜在系统图、稀疏原子干预靶点和干预强度,再通过SCM解码器传播其效果。模型完全基于由生物激励的图和表达模拟器生成的先验预测合成情节进行训练。我们在真实单细胞扰动数据和合成基准上评估PerturbPFN,涵盖效果预测、靶点识别和调控结构发现。结果表明,PerturbPFN为专业基线提供了互补权衡,以低推理成本实现有竞争力的扰动预测,同时揭示靶点、强度和系统结构的可解释中间估计。
英文摘要
Predicting cellular responses to unseen chemical perturbations is challenging due to unknown targets and mechanisms, high-dimensional expression responses, and limited experimental coverage of the large small-molecule design space. We propose PerturbPFN, a PFN-style amortized model for unknown-target perturbation prediction under a hierarchical synthetic structural prior. Instead of directly regressing high-dimensional expression responses, PerturbPFN infers a latent system graph, sparse atomic intervention targets, and intervention strengths, then propagates their effects through an SCM decoder. The model is trained entirely on prior-predictive synthetic episodes generated from biologically motivated graph and expression simulators, enabling structured in-context learning without test-time gradient updates. We evaluate PerturbPFN on both real single-cell perturbation data and synthetic benchmarks, covering effect prediction, target identification, and regulatory structure discovery. Our results show that PerturbPFN offers a complementary trade-off to specialized baselines, achieving competitive perturbation prediction with low inference cost while exposing interpretable intermediate estimates of targets, strengths, and system structure.
Comments19 pages. Accepted at the 2nd ICML Workshop on Foundation Models for Structured Data (FMSD 2026), Seoul, South Korea