ScreenShot:用于少样本组合药物筛选的基础模型
ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening
- Computational Oncology(计算肿瘤学)
- Memorial Sloan Kettering Cancer Center(纪念斯隆-凯特琳癌症中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出用于少样本组合药物筛选的基础模型ScreenShot,其经多药物筛选数据集预训练,可通过上下文学习直接预测患者样本的组合疗法响应,在预留数据集上优于基线,还可驱动主动学习策略以降低实验预算。
AI中文摘要:
采用药物组合治疗患者可降低对单一药物产生耐药性的风险。然而,由于搜索空间庞大,组合筛选成本高昂、耗时且常存在技术上的不可行性,因此难以找到有效的药物组合。预测模型可填补这一空白,但现有方法通常需要对每个样本进行分子表征,并针对每个队列进行训练,在时间和组织样本稀缺时适用性受限。为解决这一挑战,我们提出ScreenShot,这是一种在涵盖3700种药物和6000个生物样本的40个药物筛选数据集上预训练的分层Transformer模型,其架构与筛选数据的嵌套结构相匹配。给定新患者的少样本观测上下文,ScreenShot通过上下文学习预测样本对组合疗法的响应,直接基于功能测量值操作,无需微调且无需分子表征。在四个预留数据集上,ScreenShot在预测准确性和选择性有效治疗方案识别方面均优于所有基线模型。ScreenShot的内部表征可直接用于实验设计:我们利用其驱动加权k-means++主动学习策略,选择要开展的实验,以三分之一的预算实现与均匀筛选相同的命中检测效果。源代码和交互式仪表盘可通过此URL获取。
英文摘要:
Treating patients with combinations of drugs reduces the risk of resistance to any individual drug. Finding effective combinations is difficult because the large search space makes combinatorial screens prohibitively expensive, time consuming, and often technically infeasible. Predictive models can fill this gap, yet existing methods typically require molecular profiling of each sample and per-cohort training, limiting their applicability when time and tissue are scarce. To address this challenge, we introduce ScreenShot, a hierarchical transformer pretrained on 40 drug screening datasets covering 3,700 drugs and 6,000 biological samples, whose architecture mirrors the nested structure of screening data. Given a few-shot context of observations from a new patient, ScreenShot predicts the response of the sample to combination therapies through in-context learning, operating directly on functional measurements with no fine-tuning and no molecular profiling. On four held-out datasets, ScreenShot outperforms all baselines in both prediction accuracy and identification of selectively effective treatments. ScreenShot's internal representations are directly useful for experimental design: we use them to drive a weighted k-means++ active learning strategy that selects which experiments to run, achieving the same hit detection as uniform screening with a third of the budget. Source code and interactive dashboard: https://github.com/tansey-lab/screenshot.