arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过直接预测与流匹配的高效图生成

Efficient Graph Generation via Direct Prediction and Flow Matching

Susie Lu

arXiv 2610.05397首次发表:更新:

发表机构

MIT(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DiGFM,一种结合直接图预测与连续流匹配的图变换器模型,通过多步预测干净图并转换为速率向量,实现高效图生成,仅需扩散模型2.5%-15.6%的步骤,推理加速5.3-257倍,并在基准上达到或超越最先进性能。

AI 中文摘要

图结构数据的生成建模对于从药物发现到社交网络模拟等任务至关重要。在这些模型中,去噪扩散模型通过逐步学习逆转向原始图添加噪声的过程,在图生成方面取得了巨大成功。然而,扩散模型的标准噪声预测方法对于图数据而言并非最优。图生成模型的目标是学习干净图的拓扑属性,如连通性和度分布。由于预测噪声的扩散模型并未显式学习这些拓扑属性,因此模型难以输出具有所需结构统计特征的图。为解决这一挑战,我们引入了直接图流匹配(DiGFM),一种由两个目标引导的新型图变换器模型:预测干净图和提高采样效率。与主流扩散方法不同,DiGFM采用连续流匹配范式并整合直接图预测。具体而言,DiGFM通过多步过程将先验噪声分布映射到干净图分布:模型反复预测底层干净图,并利用变换将模型输出转换为指向干净图分布方向的速率向量。这种设计使DiGFM仅需基于扩散模型所需步骤的2.5%至15.6%即可生成高质量样本,从而实现5.3倍至257倍的墙钟推理时间加速。实验表明,DiGFM在通用图基准和分子数据集上优于或匹配先前最先进的模型,以显著更快的推理速度生成高度符合真实结构统计特征的图。

英文摘要

Generative modeling of graph-structured data is crucial for tasks ranging from drug discovery to social network simulation. Among these models, denoising diffusion models have achieved great success in graph generation by learning to progressively reverse a process that adds noise to the original graph. However, the standard noise-prediction approach of diffusion models is suboptimal for graph data. The goal for a graph generative model is to learn the clean graphs' topological properties, such as connectivity and degree distribution. Because a diffusion model that predicts noise does not explicitly learn these topological properties, it is challenging for the model to output graphs with the desired structural statistics. To address this challenge, we introduce Direct Graph Flow Matching (DiGFM), a novel graph transformer model guided by two goals: predict clean graphs and improve sampling efficiency. Distinct from the prevailing diffusion approach, DiGFM employs a continuous flow-matching paradigm and integrates direct graph prediction. Specifically, DiGFM maps the prior noise distribution to the clean graph distribution via a multi-step process: the model repeatedly predicts the underlying clean graph, and a transformation is employed to convert the model output to the velocity vector that points in the direction toward the clean graph distribution. This design enables DiGFM to generate high-quality samples using only 2.5% to 15.6% of the steps required by diffusion-based models, which leads to a 5.3x to 257x speedup in wall-clock inference time. Experiments demonstrate that DiGFM outperforms or matches prior state-of-the-art models across general graph benchmarks and molecular datasets, generating graphs with strong adherence to ground-truth structural statistics at significantly faster inference speeds.

CommentsNeurIPS 2026 Workshop on Geometric Distributional Deep Learning (Oral)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑