PGFS++:合成与多样性约束下的分子性质优化
PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints
浏览论文内容
中文总结 AI 辅助
PGFS++是一种感知合成的强化学习框架,通过优化反应模板与库存构建模块的应用,在改进分子目标性质的同时保留输出多样性,解决了PGFS+的奖励作弊问题。
中文摘要 AI 辅助
在早期药物发现中,优化药物相似性或结合亲和力等分子性质是一项反复开展的任务。然而,在无约束化学空间中优化得到的分子若无法合成,实际应用价值有限。Policy Gradient for Forward Synthesis(PGFS)是一种用于分子优化的感知合成强化学习方法,但它采用反应物嵌入预测导致反应物选择间接,我们的研究表明这会限制学习效果。我们首先开发了PGFS+,其中反应模板和第二反应物由可训练的嵌入查找表表示,结合更有效的评分函数和强化学习算法,PGFS+显著提升了目标性质。不过,它存在奖励作弊的失效模式:强大的反应物搜索可将不同输入分子映射到同一高奖励“磁体分子”,在提升奖励的同时导致输出多样性崩溃。因此,我们提出PGFS++,这是一种针对输入特定分子优化的感知合成强化学习框架。给定输入分子,PGFS++将其视为正向合成轨迹的起点,应用学习到的反应模板与兼容的库存构建模块,生成具有改进目标性质、明确合成路线且与输入结构相似的分子。分子优化任务实验表明,PGFS++在提升目标性质的同时保持了高输出多样性。
英文摘要
Improving molecular properties, such as drug-likeness or binding affinity, is a recurring task in early-stage drug discovery. However, molecules optimized in an unconstrained chemical space have limited practical value if they cannot be synthesized. Policy Gradient for Forward Synthesis (PGFS) is a synthesis-aware reinforcement learning method for molecular improvement, but its use of reactant embedding prediction makes reactant selection indirect, which, as we show, limits learning effectiveness. We first develop PGFS+, in which reaction templates and second reactants are represented by trainable embedding lookup tables. Combined with a more effective scoring function and RL algorithm, PGFS+ significantly improves the desired property. However, it exposes a reward-hacking failure mode: a powerful reactant search can map diverse input molecules to the same high-reward magnet molecule, improving the reward while collapsing the output diversity. We therefore introduce PGFS++, a synthesis-aware reinforcement learning framework for input-specific molecular improvement. Given an input molecule, PGFS++ treats it as the start of a forward-synthesis trajectory, applies learned reaction templates with compatible in-stock building blocks, and produces a molecule with improved target properties, an explicit synthesis route, and structural similarity to the input. Experiments on molecular improvement tasks show that PGFS++ improves target properties while preserving high output diversity.
发表机构
- Graphcore(格洛科普(Graphcore))
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。