AI 中文总结
研究针对生成性分子模型评估指标不足的问题,引入受工业工作流程启发的六阶段筛选基准刺猬(HEDGEHOG),在标准化协议下评估三个模型类别的23个分子生成器,揭示了当前分子生成器难以同时满足多方面筛选的局限性。
AI 中文摘要
生成性分子模型可通过从头提出新的候选化合物来支持早期药物发现。但常用的评估分子生成器的指标难以反映生成的化合物在医学上是否合理及是否适用于下游计算,易产生模型评估误报等问题。我们引入了刺猬(HEDGEHOG),这是一个统一的六阶段筛选基准,受工业命中识别工作流程启发。我们在标准化协议下评估了三个模型类别的23个分子生成器。在230,000个生成的分子中,只有0.65%的初始分子能通过所有阶段。结果揭示了当前分子生成器的核心局限性:在单一标准下看似可接受的分子很少能同时满足药物化学、合成、对接和3D构象筛选。
英文摘要
Generative molecular models can support early drug discovery by proposing new candidate compounds de novo. In practice, useful candidates must balance target-relevant activity, physicochemical properties, and other multiparameter design constraints. However, standard metrics commonly used to evaluate molecular generators only weakly reflect whether the generated compounds are medicinally plausible and suitable for downstream computation. This can produce an incomplete view of model performance and inefficient use of computational resources. We introduce HEDGEHOG, a unified six-stage filtration benchmark that is constructed as a hit identification workflow: (i) preprocessing; (ii) physicochemical descriptor screening; (iii) structural alerts and graph-sanity checks; (iv) synthesis feasibility; (v) docking; and (vi) three-dimensional pose and interaction checks. We evaluated 22 generative models in a KRAS G12D case study, using three runs of 1,000 requested generation attempts per model. The models showed different patterns of attrition, and final survival ranged from 0 to 127 molecules per run. None of the standard metrics showed a significant association with final survival after correction for multiple testing. HEDGEHOG provides a reproducible benchmark for evaluating molecular generative models by measuring molecule survival through chemical filters. The framework identifies stage-wise failure modes across generator classes and provides a practical basis for developing molecular generators better aligned with early drug discovery.
Comments26 pages (including References and Appendix sections), 33 tables, 6 figures, 1 supplementary file