发表机构
AMD Silo AI; AstraZeneca(AMD Silo AI; 阿斯利康)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出开放模块化智能体MAGI,用于协调分子设计工具,在真实药物发现项目中,其性能受限于预测模型的适用性,而非生成路径。
AI 中文摘要
智能体系统越来越多地协调分子设计工具,但尚不清楚堆栈中的哪一层限制了真实项目的结果。我们开发了MAGI,一个开放模块化智能体,它能够编写目标、启动并监控优化、解释构效关系,并相应调整其策略。MAGI可以直接通过LLM生成分子,或委托给REINVENT 4,评分服务在通用契约下可互换。我们在来自三家制药公司的九个回顾性先导化合物优化活动中对其进行了测试,并在固定的时间截点下重放。两种途径都产生了有效的结构:LLM提出的方案更接近局部化学,并在更少的操作中达到相当或更高的主要活性,而REINVENT探索了更广泛的化学空间。一个活动是否达到其目标取决于预测模型,而非生成途径:达标率随所提出化学的模型准确性而变化,一旦该化学超出模型适用域,达标率即下降。另外,一项盲评评估了MAGI的输出是否能通过专家审查:化学家无法区分智能体提出的方案与保留化合物,并认为SAR推理大致合理但不完整。综合这些结果,MAGI定位为可插入现有计算化学工作流的协调层。然而,真实项目的上限目前仍由评分器的适用性决定,而非工具编排。
英文摘要
Agentic systems increasingly coordinate molecular-design tools, but it is unclear which layer of the stack limits outcomes on real projects. We developed MAGI, an open modular agent that authors objectives, launches and monitors optimization, interprets structure--activity relationships, and revises its strategy accordingly. MAGI generates molecules either directly through the LLM or by delegating to REINVENT 4, with scoring services interchangeable behind a common contract. We tested it across nine retrospective lead-optimization campaigns from three pharmaceutical companies, replayed under fixed temporal cutoffs. Both routes produced valid structures: LLM proposals stayed closer to local chemistry and reached comparable or higher primary activity in fewer operations, whereas REINVENT explored broader chemical space. Whether a campaign met its objective depended on the predictive models, not on the generation route: attainment followed model accuracy on the chemistry proposed, dropping once that chemistry moved outside the model's applicability domain. Separately, a blinded evaluation asked whether the MAGI's output could pass as expert work: chemists were not able to discriminate agentic proposals from held-out compounds, and judged the SAR reasoning broadly plausible yet incomplete. Together, these results position MAGI as a coordination layer pluggable into existing computational chemistry workflows. The ceiling on real projects, however, remains currently set by scorer applicability rather than by tool orchestration.