arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05431cs.AI

PharmAgent:基于冻结语言模型的约束感知搜索用于分子优化

PharmAgent: Constraint-Aware Search with Frozen Language Models for Molecular Optimization

  • City University of Hong Kong (Dongguan)(香港城市大学(东莞))
  • The Hong Kong Polytechnic University(香港理工大学)
  • Tencent AI Lab(腾讯AI实验室)
  • Zhengzhou University(郑州大学)
  • Northwestern Polytechnical University(西北工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Nihui Shao, Guanxing Chen, Jilong Shi, Zhengyang Bai, Haohuai He, Zhenchao Tang, Qiujie Lv, Yu-An Huang, Zhi-An Huang

AI总结:

PharmAgent提出一种基于冻结语言模型和自适应外部状态的约束感知分子搜索方法,通过拉格朗日控制器和结构索引重放优化分子,在多个任务上显著超越现有基线。

AI中文摘要:

分子优化必须在有限的评估预算内提高目标活性并满足可开发性约束。经典方法需要定制的规则或训练来纳入化学指令和性质反馈。冻结语言模型可以根据这些信息调节编辑,但需要显式的约束控制和相关经验。因此,我们提出了PharmAgent,一种由自适应外部状态驱动的约束感知分子搜索方法。其拉格朗日控制器将接受状态中的违规转化为累积的约束压力,并将此历史与当前性质测量分开存储。结构索引重放通过相关的评估转换补充此反馈,以指导后续提议。随着课程逐步激活约束,候选者和当前最优解在同一当前目标下进行比较,接受状态决定下一次乘数更新。我们推导出一个精确恒等式,刻画了接受状态违规如何在控制器乘数中累积。在五个任务和五次独立运行中,PharmAgent在仅目标搜索中实现了目标得分曲线下面积总和(AUC)为3.9208,比MOLLEO提高了37.3%。在在线约束下,其实现了性质调整后的AUC为0.7076,比最强在线基线ExLLM提高了53.8%。这些结果在仅目标和约束感知搜索中均在所有评估方法中排名第一。在线比较覆盖了所有五个基线框架。完整系统在目标质量、性质调整性能和帕累托超体积方面领先于所有消融变体。所有五个分子案例均达到可行最终状态,记录了目标增益和权衡。

英文摘要:

Molecular optimization must improve target activity and satisfy developability constraints within limited evaluation budgets. Classical methods require tailored rules or training to incorporate chemical instructions and property feedback. Frozen language models can condition edits on this information, but need explicit constraint control and relevant experience. We therefore present PharmAgent, a constraint-aware molecular search method driven by adaptive external state. Its Lagrangian controller translates violations in accepted states into accumulated constraint pressure, keeping this history separate from current property measurements. Structure-indexed replay complements this feedback with relevant evaluated transitions that guide subsequent proposals. As a curriculum progressively activates constraints, candidates and the incumbent are compared under the same current objective, and the accepted state determines the next multiplier update. We derive an exact identity that characterizes how accepted-state violations accumulate in the controller's multipliers. Across five tasks with five independent runs, PharmAgent achieves a summed area under the target-score curves (AUC) of 3.9208 in target-only search, improving over MOLLEO by 37.3%. With online constraints, it achieves a property-adjusted AUC of 0.7076, improving over the strongest online baseline, ExLLM, by 53.8%. These results rank first among all evaluated methods in both target-only and constraint-aware search. The online comparison covers all five baseline frameworks. The full system leads every ablation variant in target quality, property-adjusted performance, and Pareto hypervolume. All five molecular cases reach feasible final states, documenting target gains and trade-offs.

补充信息

↑