arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

闭环生成选择:收敛性、记忆与噪声预言机

Closed-Loop Generative Selection: Convergence, Memory, and Noisy Oracles

Konstantin Fackeldey, Christof Schütte

arXiv 2607.22211首次发表:更新:

AI 中文总结

研究闭环生成选择算法,通过恢复马尔可夫结构证明收敛性与运行时间界限,分析模型记忆作用,还扩展到多目标搜索和噪声预言机,得出评估最少策略并经实验证实,最后提出三个开放问题。

AI 中文摘要

闭环生成选择已成为计算药物发现的常用方法:学习生成模型提出候选分子,适应度预言机对其评分,保留最佳分子,然后在下一轮之前在这个精英集上重新训练模型。尽管广泛使用,但该方法缺乏严格的收敛理论,主要因为每轮重新训练模型破坏了经典进化算法分析所依赖的马尔可夫性质。我们为此类算法开发了一个自包含的收敛性和预期运行时间理论。通过在扩大的状态空间上恢复马尔可夫结构,我们表明精英主义使搜索具有吸收性,并证明了几乎必然收敛以及将搜索分解为逃离每个适应度水平所花费时间的运行时间界限。接着分析了模型记忆的作用——即模型基于多少过去的数据进行训练。当学习随更多数据稳步改善时,更深的记忆无害;当并非如此时,退出时间分析确定了最优记忆深度,并表明过多记忆实际上会减缓收敛。该理论扩展到多目标搜索和噪声预言机:我们量化了在轻尾噪声下多少次重复评估能证明进展,以及稳健估计器在重尾情况下如何恢复保证。根据预言机评估(药物设计中的真正瓶颈)重新表述,分析得出了一个具体的、评估最少的策略。一项可重复研究证实了这些预测,包括过多记忆的惊人成本。最后我们提出了三个开放问题。

英文摘要

Closed-loop generative selection has become a workhorse of computational drug discovery: a learned generative model proposes candidate molecules, a fitness oracle scores them, the best are kept, and the model is retrained on this elite set before the next round. Despite its wide use, the method has lacked a rigorous convergence theory, largely because retraining the model each round breaks the Markov property on which classical evolutionary-algorithm analysis relies. We develop a self-contained theory of convergence and expected running time for this class of algorithms. By recovering a Markov structure on an enlarged state space, we show that elitism makes the search absorbing, and we prove almost-sure convergence together with a runtime bound that decomposes the search into the time spent escaping each fitness level. We then analyse the role of the model's memory---how much of the past it is trained on. When learning improves steadily with more data, deeper memory never hurts; when it does not, an exit-time analysis pinpoints the optimal memory depth and shows that excess memory can actually slow convergence. The theory extends to multi-objective search and to noisy oracles: we quantify how many repeated evaluations certify progress under light-tailed noise, and how robust estimators restore guarantees under heavy tails. Recast in terms of oracle evaluations - the true bottleneck in drug design - the analysis yields a concrete, evaluation-minimal strategy. Areproducible study confirms the predictions, including the surprising cost of excess memory. We close with three open problems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑