PertMind:通过细胞扰动数据的强化学习在大语言模型中引出涌现生物推理能力
PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data
浏览论文内容
中文总结 AI 辅助
本研究提出PertMind模型,结合可信轨迹监督初始化与多层面强化信号,以细胞扰动图谱为强化学习环境训练,实现生物推理能力的涌现,且可迁移至多类生物任务。
中文摘要 AI 辅助
大语言模型能够描述机制,但其可扩展的后训练仍依赖于成本高昂、人工整理的生物推理轨迹。本文表明,细胞扰动图谱可替代成为强化学习环境,其中测得的基因响应为生物推理提供可计算的奖励。我们提出了PertMind,它结合了可信轨迹的监督初始化以及基因、通路和格式层面的强化信号。仅在正向扰动-响应预测上进行训练的PertMind,在保留通用语言能力的同时,提升了未见过的细胞环境中的响应推理能力;且无需针对特定任务进行后训练,即可迁移至反向扰动识别、双扰动推理、表型筛选优先级排序及生物过程解释任务。PertMind还生成了生物图谱,在多尺度下游任务中支持具有竞争力的基因、细胞和供体表征。这些结果支持了如下假设:对实验终点的强化学习可浓缩预训练模型已具备的可复用生物策略。更广泛而言,源自扰动的强化学习提供了一条可扩展路径,将不断扩展的实验图谱转化为通用生物推理的训练环境。
英文摘要
Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning. We introduce PertMind, which combines trusted-trajectory supervised initialization with gene-, pathway-, and format-level reinforcement signals. Although trained only on forward perturbation-response prediction, PertMind improves response inference in unseen cellular contexts while retaining general language capabilities. It also transfers, without task-specific post-training, to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation. PertMind further generates biological profiles that support competitive gene, cell, and donor representations across multiscale downstream tasks. These results support the hypothesis that reinforcement on experimental endpoints can concentrate reusable biological strategies already accessible to pretrained models. More broadly, perturbation-derived reinforcement learning offers a scalable route for transforming expanding experimental atlases into training environments for general-purpose biological reasoning.