发表机构
Technical University of Denmark(丹麦技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种基于算法控制随机合成和混合物测试的信息密集合成方法,将分子发现实验复杂度从线性降至对数或常数,在模拟中比贝叶斯优化少一个数量级实验。
AI 中文摘要
机器学习可以通过设计分子和规划实验来加速分子发现。然而,许多科学挑战需要具有非常稀有性质的分子,在这种稀疏设置下,现有算法相比随机猜测几乎没有增益。我们提出了一种利用算法控制的随机合成来高效搜索分子空间大区域的方法。我们不是设计、制造和测试单个分子,而是设计和制造复杂混合物,将其作为一个整体进行测试,然后解卷积分子-活性图谱。我们优化合成以编码最大信息。理论上,这种方法可以将从$d$个候选物中找到最优分子所需的实验次数从$\mathcal{O}(d)$减少到$\mathcal{O}(\log d)$或$\mathcal{O}(1)$。在模拟中,在估计的蛋白质适应性景观上,它比现有贝叶斯优化方法少一个数量级的实验就能找到活性分子。
英文摘要
Machine learning can accelerate molecular discovery by designing molecules and planning experiments. However, many scientific challenges demand molecules with very rare properties, and in this sparse setting, existing algorithms offer little gain over random guessing. We propose a method to efficiently search large regions of molecular space using algorithmically controlled stochastic synthesis. Rather than design, make and test individual molecules, we design and make complex mixtures, test them as a pool, then deconvolute the molecule-activity map. We optimize synthesis to encode maximal information. Theoretically, this approach can reduce the number of experiments required to find the optimal molecule among $d$ candidates from $\mathcal{O}(d)$ to $\mathcal{O}(\log d)$ or $\mathcal{O}(1)$. In simulation, on estimated protein fitness landscapes, it finds active molecules with an order of magnitude fewer experiments than existing Bayesian optimization methods.