发表机构
Fudan University; Shanghai Innovation Institute; Soochow University; Southeast University; East China Normal University(复旦大学; 上海创新研究院; 苏州大学; 东南大学; 华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MARCO提出了一种基于评估器的多轮强化学习框架,通过提议-反馈-修订轨迹训练分子编辑器,在MuMOInstruct基准上实现了性质成功率和相似性的最佳平衡。
AI 中文摘要
分子优化本质上是迭代的:提出一个候选分子,根据多个目标进行评估,并在保持与源分子关系的同时进行修改。大多数遵循指令的模型反而只输出一个编辑后的分子,迫使有效性、性质改进和相似性控制都集中在一个响应中。我们引入了MARCO,一个基于评估器的强化学习框架,该框架在受限的提议-反馈-修订轨迹上训练分子编辑器。MARCO将分阶段的回合奖励聚合为无折扣的轨迹回报,用于组相对策略优化。我们评估了这种训练的两种结果:Same-1在单响应预算下测试训练后的策略,而Same-5则测试同一策略在最多五个响应可用时是否能利用验证器的反馈。在三个目标的MuMOInstruct基准、三个Qwen骨干模型以及已见/未见指令划分中,SFT初始化的MARCO在每个报告的主要设置中都获得了性质成功率和相似性的最高乘积。Same-5在测试预算下进一步提高了观察到的得分,而四目标和公共检查点实验则测试了跨约束集和初始化机制的迁移能力。
英文摘要
Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response. We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal--feedback--revision trajectories. MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optimization. We evaluate two consequences of this training: Same-1 tests the trained policy under a one-response budget, while Same-5 tests whether the same policy can use verifier feedback when up to five responses are available. Across the three-objective MuMOInstruct benchmark, three Qwen backbones, and seen/unseen instruction splits, SFT-initialized MARCO obtains the highest product of property success rate and similarity in every reported primary setting. Same-5 further improves the observed score under the tested budget, while four-objective and public-checkpoint experiments test transfer across constraint sets and initialization regimes.