发表机构
Biogen(渤健公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EW-SFT是一种统一有效的分子优化方法,通过奖励指导精英选择并基于预训练损失更新模型,适用于多种分子生成器与设计任务,性能优于原生优化器。
AI 中文摘要
目标导向优化对于引导分子生成器提出具有所需特性的候选分子至关重要。然而,该优化常通过策略梯度强化学习实现,其需要生成轨迹对数概率,形式依赖于模型架构和生成过程,这使得优化器难以跨架构和条件生成设计复用。监督微调无需此类机制,但其更新由固定数据集驱动,奖励从未参与更新。我们提出精英加权监督微调(Elite-Weighted Supervised Fine-tuning,EW-SFT),该方法利用奖励指导高得分分子的精英选择,并基于该集合中分子的预训练损失更新模型。消融实验表明,奖励信息主要通过精英选择传递,而非所选集合内的连续加权。由于更新仅使用已得分分子和模型自身损失,该规则适用于自回归、掩码扩散、离散流生成器,以及从头设计、基序扩展、连接子设计任务。在对两种激酶参考化合物使用3D形状对齐神谕的固定预算下,EW-SFT始终优于对应原生优化器;在对四种保留参考使用2D相似度神谕的目标导向优化中,其进一步提升性能,且在无轨迹级RL公式的样本效率基准上达到可比性能。这些结果表明,EW-SFT是跨分子生成器、设计约束、参考和神谕的统一且有效的优化器。
英文摘要
Goal-directed optimization is essential for steering molecular generators to propose candidates with desired properties. However, it is often implemented with policy-gradient reinforcement learning, which requires a generation-trajectory log-probability whose form depends on the model architecture and generation procedure. This makes an optimizer difficult to reuse across architectures and conditional generative designs. Supervised fine-tuning needs none of that machinery, but its update is driven by a fixed dataset, so the reward never enters the update. We introduce Elite-Weighted Supervised Fine-tuning (EW-SFT), which uses reward to guide elite selection of high-scoring molecules, and updates the model by its own pretraining loss on that set. Ablations show that reward information is passed primarily through elite selection, rather than through continuous weighting within the selected set. Because the update consumes only scored molecules and the model's native loss, the same rule applies across autoregressive, masked-diffusion, and discrete-flow generators, and across de novo, motif-extension, and linker-design tasks. Under a fixed budget of 3D shape alignment oracle calls on two kinase reference compounds, EW-SFT consistently outperforms the corresponding native optimizers. It further improves goal-directed optimization under a 2D similarity oracle on four held-out references and achieves comparable performance on a sample-efficiency benchmark without a trajectory-level RL formulation. These results demonstrate that EW-SFT is a unified and effective optimizer across molecular generators, design constraints, references, and oracles.