GraphVQ:基于上下文量化图令牌的结构感知自回归解码
GraphVQ: Structure-Aware Autoregressive Decoding over Context-Quantized Graph Tokens
浏览论文内容
中文总结 AI 辅助
GraphVQ 通过上下文量化的图令牌和成对条件解码,解决自回归图生成中的结构障碍,显著提升环生成和分布忠实度,在多个数据集上取得最优性能。
中文摘要 AI 辅助
图基础模型需要离散的令牌表示,但将图转化为可生成的令牌序列面临一个结构性障碍:跨越序列化窗口的边无法在一次生成中输出,因此单次自回归生成器会系统性地欠生成环;同时,单一的全局条件无法区分候选边。GraphVQ 消除了这两个障碍:节点上下文(特征加上多阶广度优先序列化下的局部边掩码)通过一个具有 BCE 校准的伯努利边解码的 VQ-VAE 量化到共享码本中,第二阶段的架构感知解码器基于令牌导出的成对特征生成全局邻接矩阵,其相对于任何全局摘要条件的必要性在限定范围的不可行性结果中被正式化。分词器以 0.86–0.99 的准确率重建节点特征,并以 AUROC >= 0.89(ECE <= 0.007)解码局部边。在四个数据集上使用相同的分割协议,成对条件将 PROTEINS 上的轨道 MMD 从 0.248 改善到 0.174,在环压力测试上改善 3.4 倍,并在随机标签控制上消失——这是属性-拓扑耦合的特征——因此增益恰好归因于属性携带边相关信号的情况。GraphVQ 在 PROTEINS 上的学习生成器中排名第一,在 SYN-COMM 上并列第一,并在三个数据集上将轨道 MMD 比单阶段生成改善 2.7–17 倍,种子级自助法区间确认排名不是种子噪声;在 MUTAG 上,未加权边目标欠生成,并如实报告。这些结果将自回归图生成的结构控制定位于条件的粒度:成对级令牌上下文将量化词汇转化为分布忠实图生成和未来令牌级预训练的可利用容量轴。
英文摘要
Graph foundation models need a discrete token representation, but casting a graph as a generatable token sequence faces a structural obstacle: edges spanning beyond the serialization window cannot be emitted in one pass--so one-pass autoregressive generators systematically under-produce cycles--and a single global condition cannot tell candidate edges apart. GraphVQ removes both obstacles: node contexts--features plus a local edge mask under multi-order breadth-first serialization--are quantized into a shared codebook by a VQ-VAE with BCE-calibrated Bernoulli edge decoding, and a second-stage structure-aware decoder emits the global adjacency conditioned on token-derived pair features, whose necessity over any global-summary condition is formalized in a scoped impossibility result. The tokenizer reconstructs node features at 0.86--0.99 accuracy and decodes local edges at AUROC >= 0.89 (ECE <= 0.007). Under one same-split protocol on four datasets, pair conditioning improves orbit MMD 0.248 -> 0.174 on PROTEINS and 3.4x on a ring stress test, and vanishes on a random-label control--the signature of attribute--topology coupling--so the gain is claimed exactly where attributes carry edge-relevant signal. GraphVQ ranks first among learned generators on PROTEINS, ties for first on SYN-COMM, and improves orbit MMD 2.7--17x over one-stage generation on three datasets, with seed-level bootstrap intervals confirming the rankings are not seed noise; on MUTAG the unweighted edge target under-generates and is reported as such. These results locate the structural control of autoregressive graph generation in the granularity of the condition: pair-level token context turns a quantized vocabulary into a usable capacity axis for distribution-faithful graph generation and future token-level pretraining.
发表机构
- China Life Insurance Company Ltd.(中国人寿保险股份有限公司)
- Beijing Institute of Technology(北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。