arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16419cs.LGcs.AIq-bio.QM

PertMind:通过细胞扰动数据的强化学习在大语言模型中引出涌现生物推理能力

PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

Zhenchao Tang, Xiaogang Xu, Jiafei Wu, Jiahui Guan, Bo Li, Tianxu Lv, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Xun Lin, Zhipeng Deng, Zhaoxi… 展开作者

Zhenchao Tang, Xiaogang Xu, Jiafei Wu, Jiahui Guan, Bo Li, Tianxu Lv, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Xun Lin, Zhipeng Deng, Zhaoxing Li, Guanxing Chen, Yaokun Li, Mengran Li, Songming Zhang, Zhe Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出PertMind模型,结合可信轨迹监督初始化与多层面强化信号,以细胞扰动图谱为强化学习环境训练,实现生物推理能力的涌现,且可迁移至多类生物任务。

中文摘要 AI 辅助

大语言模型能够描述机制,但其可扩展的后训练仍依赖于成本高昂、人工整理的生物推理轨迹。本文表明,细胞扰动图谱可替代成为强化学习环境,其中测得的基因响应为生物推理提供可计算的奖励。我们提出了PertMind,它结合了可信轨迹的监督初始化以及基因、通路和格式层面的强化信号。仅在正向扰动-响应预测上进行训练的PertMind,在保留通用语言能力的同时,提升了未见过的细胞环境中的响应推理能力;且无需针对特定任务进行后训练,即可迁移至反向扰动识别、双扰动推理、表型筛选优先级排序及生物过程解释任务。PertMind还生成了生物图谱,在多尺度下游任务中支持具有竞争力的基因、细胞和供体表征。这些结果支持了如下假设:对实验终点的强化学习可浓缩预训练模型已具备的可复用生物策略。更广泛而言,源自扰动的强化学习提供了一条可扩展路径,将不断扩展的实验图谱转化为通用生物推理的训练环境。

英文摘要

Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning. We introduce PertMind, which combines trusted-trajectory supervised initialization with gene-, pathway-, and format-level reinforcement signals. Although trained only on forward perturbation-response prediction, PertMind improves response inference in unseen cellular contexts while retaining general language capabilities. It also transfers, without task-specific post-training, to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation. PertMind further generates biological profiles that support competitive gene, cell, and donor representations across multiscale downstream tasks. These results support the hypothesis that reinforcement on experimental endpoints can concentrate reusable biological strategies already accessible to pretrained models. More broadly, perturbation-derived reinforcement learning offers a scalable route for transforming expanding experimental atlases into training environments for general-purpose biological reasoning.

补充信息

↑