发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出反思型智能体EASER,通过低秩矩阵属性接口连接文献证据与扩散生成,以显式假设制定实现多目标抗菌肽设计,取得最优多目标性能。
AI 中文摘要
大型语言模型(LLMs)能够对科学文献进行推理以制定设计策略,但无法可靠地将其应用于生物序列的实现。虽然蛋白质生成模型学习了序列模式,但它们缺乏整合文献证据进行多步反思性推理的能力,从而在科学推理与序列操作之间形成了证据到执行的鸿沟。我们提出了EASER(具有反思能力的证据感知序列工程),一个通过离线训练、固定的低秩矩阵学习到的属性接口,将推理与序列生成连接起来的反思型智能体。该智能体通过组合这些矩阵来引导扩散生成器,提出基于检索到的证据、序列上下文和过往结果的干预假设(锚点、可编辑位置、控制系数)。一种“探测与引导”机制验证干预措施,并根据预测的属性响应分配样本,结果反思为后续决策提供信息。在多目标抗菌肽设计(优化活性、非溶血性和非毒性)的评估中,在相同的决策条件下,显式假设制定比直接动作生成带来了更好的多目标性能。消融研究验证了证据检索、情节历史、反思和“探测与引导”的重要性。在六次重复试验中,与竞争基线相比,EASER在筛选候选物上获得了最高的平均超体积和最低的平均IGD+。我们的工作展示了可执行的属性接口和迭代反馈如何将科学推理与目标肽序列生成联系起来。
英文摘要
Large language models (LLMs) can reason over scientific literature to devise design strategies, yet fail to reliably implement them for biological sequences. While protein generative models learn sequence patterns, they lack the capacity to incorporate literature evidence for multi-step reflective reasoning, forming an evidence-to-execution gap between scientific reasoning and sequence manipulation. We present EASER (Evidence-Aware Sequence Engineering with Reflection), a reflective agent bridging reasoning and sequence generation via a learned property interface of offline-trained, fixed low-rank matrices. The agent steers a diffusion generator by combining these matrices, proposing intervention hypotheses (anchors, editable positions, control coefficients) grounded in retrieved evidence, sequence context and past results. A Probe-and-Steer mechanism validates interventions and allocates samples according to predicted property responses, with outcome reflection informing subsequent decisions. Evaluated on multi-objective antimicrobial peptide design (optimizing activity, non-hemolysis and non-toxicity), explicit hypothesis formulation delivers better multi-objective performance than direct action generation under identical decision conditions. Ablation studies verify the importance of evidence retrieval, episodic history, reflection and Probe-and-Steer. Over six repeated trials, EASER obtains the highest mean hypervolume and lowest mean IGD+ on screened candidates compared with competing baselines. Our work demonstrates how an executable property interface and iterative feedback link scientific reasoning to targeted peptide sequence generation.
Comments24 pages, 8 figures