AI 中文总结
研究利用大语言模型进行分子生成时面临的问题,提出LLMol框架,采用监督学习与强化学习结合的两阶段训练范式,引入RLVR及GRPO算法,有效处理多种分子设计任务,实验证明其性能优于现有方法。
AI 中文摘要
利用大语言模型进行分子生成在化学和药物设计中显示出巨大潜力。当前方法主要依赖监督训练或有限数据集微调,不足以捕捉复杂分子设计目标。我们提出LLMol,一个有原则的强化学习框架,直接将可验证奖励纳入目标分子生成。它将分子设计作为目标条件序列预测任务,采用监督学习和强化学习两阶段训练范式。在第一阶段,对大语言模型进行监督微调以捕捉化学语法和分子分布;第二阶段引入带可验证奖励的强化学习(RLVR),并采用组相对策略优化(GRPO)来解决离散序列优化中的高方差和不稳定性问题。实验结果表明LLMol在各种分子基准测试中始终优于现有方法。
英文摘要
Leveraging large language models (LLMs) for molecular generation has shown remarkable potential in chemical and drug design. Current methods primarily rely on supervised training or fine-tuning with limited datasets, which are insufficient to capture complex molecular design objectives. While some approaches attempt to guide generation toward specific goals, they often lack direct optimization mechanisms, making it difficult to align generated molecules with desired properties. To tackle these challenges, we propose \textbf{LLMol}, a principled reinforcement learning framework that directly incorporates verifiable rewards for targeted molecule generation. The key insight is to formulate molecular design as a goal-conditioned sequence prediction task, where verifiable rewards serve as explicit supervision to drive generation toward desired objectives. LLMol follows a two-stage training paradigm combining supervised learning and reinforcement learning. In the first stage, large language models are supervised fine-tuned to capture chemical syntax and molecular distributions. In the second stage, we introduce Reinforcement Learning with Verifiable Rewards (RLVR), which directly integrates property-based reward signals to guide molecular generation toward task-specific objectives. To address the high variance and instability common in discrete sequence optimization, we adopt Group Relative Policy Optimization (GRPO), a stable on-policy algorithm that smooths reward signals and improves training robustness. This framework enables LLMol to effectively handle a range of molecular design tasks, including single-property targeting (e.g., penalized logP, QED) and structure-constrained optimization. Experimental results demonstrate that LLMol consistently outperforms existing methods, achieving higher success rates and improved efficiency across diverse molecular benchmarks.
Comments13 pages, 4 figures