发表机构
University of Antwerp(安特卫普大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AutoGrable是一种无需训练图模型的表格到图构造方法,通过标签对齐风险打分选择最优列子集,在受控和真实任务中性能优于对比方法,还可在图无用时弃权构建。
AI 中文摘要
图学习以图为前提,而表格和关系型数据库并不自带图。将图神经网络(GNN)应用于表格时,需手动、通过模式启发式方法,或针对每个候选图训练模型后保留最优者,来决定哪些实体作为节点、哪些节点相连以及通过何种关系。本文提出一种无需训练图模型的准则:在最小的表格到图的抽象中,每一行对应一个节点,受1-WL约束的消息传递GNN仅将该结构视为行划分为颜色细化类别的划分;当该划分将具有不同标签的行分开且不拆分具有相同标签的行时,该结构对任务而言是良好的。AutoGrable将此准则转化为构造过程:对于关联结构,划分由所选列固定,因此构建图简化为选择列,本文通过标签对齐风险对候选子集打分——该风险是其块上最优常数预测器的保留风险,并由衡量块填充稀疏程度的占用项进行惩罚。该打分无需实例化图,也无需训练GNN,因此AutoGrable可在子集空间中以低成本贪婪搜索,为单表格和外键模式返回所得图结构。实验表明,在候选图空间中,该打分可丢弃大部分候选图同时保留最优者;在受控任务中,AutoGrable可恢复生成标签的列,在固定预测器下的真实任务中,其性能优于固定、随机和感知任务的构造方法;它也是所对比方法中唯一能在图无帮助时弃权(不执行)构建图的方法。
英文摘要
Graph learning presupposes a graph, and tables and relational databases do not come with one. Applying a GNN to them requires deciding which entities become nodes, which of them to connect, and through which relations---a decision made by hand, by schema heuristics, or by training a model on every candidate graph and keeping the best. We give a criterion that requires no trained graph model. In the minimal table-to-graph abstraction each row is a node, so a message-passing GNN, bounded by 1-WL, sees a construction only as a partition of the rows into colour-refinement classes: a construction is good for a task when that partition separates rows with different labels and does not split rows that share one. AutoGrable turns this criterion into a construction procedure. For incidence constructions the partition is fixed by the selected columns, so building a graph reduces to choosing them, and we score a candidate subset by a label-alignment risk: the held-out risk of the best predictor constant on its blocks, penalised by an occupancy term measuring how thinly the blocks are populated. The score materialises no graph and trains no GNN, so AutoGrable can search the space of subsets greedily and cheaply, and returns the resulting grable for single tables and for foreign-key schemas alike. Our experiments show that over a space of candidate graphs the score discards a large fraction while retaining the best; that AutoGrable recovers the columns that generate the label on controlled tasks and outperforms fixed, random, and task-aware constructors on real tasks under a fixed predictor; and that it is the only method compared that can decline to build a graph when none helps.
Comments28 pages, 4 figures