发表机构
École normale supérieure Paris-Saclay; Mohamed bin Zayed University of Artificial Intelligence; LIX, CNRS, École Polytechnique, IP Paris(巴黎-萨克雷高等师范学校; 穆罕默德·本·扎耶德人工智能大学; 巴黎综合理工学院,法国国家科学研究中心,LIX实验室,巴黎综合理工学院联盟)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出在预训练VAE的潜在表示上通过流匹配生成分子图,仅在最后解码,实现高质量、高效率的生成,并支持性质引导和有效性感知生成。
AI 中文摘要
现代图生成模型通常直接在离散图空间中操作,显式生成节点和边变量,随着图的增大,这种方法的成本会变得高昂。在本文中,我们直接在从预训练变分自编码器获得的整个图的潜在表示上进行生成,该自编码器具有高重建保真度。通过流匹配获得的生成表示仅在最后一步进行解码。在规模不断增大的分子基准上,我们的方法实现了强大的有效性和FCD,同时与最先进的显式图生成模型相比,提供了良好的质量-效率权衡。该公式的一个主要优点是图表示只需学习一次,之后同一个表示可以在多个生成目标中重用,无需重新训练。我们展示了由分子性质引导的生成,并进一步通过直接在潜在空间中学习的分类器引入了有效性感知的生成。所有代码将在接受后公开。
英文摘要
Modern graph generative models typically operate directly in the discrete graph space, explicitly generating node and edge variables, which can become costly as graphs grow. In this paper, we perform generation explicitly on latent representations of entire graphs obtained from a pretrained Variational Autoencoder with high reconstruction fidelity. The generated representations, obtained through flow matching, are then decoded only at the final step. Across molecular benchmarks of increasing size, our approach achieves strong validity and FCD while offering a favorable quality-efficiency trade-off compared with state-of-the-art explicit graph generative models. One of the main advantages of this formulation is that the graph representation only needs to be learned once, after which the same one can be reused across multiple generative objectives without retraining. We demonstrate generation guided by molecular properties and further introduce validity-aware generation though a classifier learned directly in latent space. All code will be made available upon acceptance.