发表机构
Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出首个无需重新训练即可泛化到新图的全归纳基数估计器FICE,通过耦合GNN组件实现,在10个KG上较基线显著降低中位数q误差并优化尾部行为。
AI 中文摘要
对知识图谱(KG)上的基本图模式(BGP)SPARQL查询进行查询优化,需要准确的基数估计。最近提出的基于学习的估计器性能优于基于统计和采样的方法,但存在一个限制,阻碍了它们在实际三元组存储中的应用:它们是转导式的,当底层图发生变化或应用于新图时需要重新训练。我们提出FICE(全归纳基数估计器,Fully Inductive Cardinality Estimation),这是首个针对知识图谱上BGP查询的基于学习的基数估计器,可泛化到完全未见过的图(包括未见过的关系),无需任何重新训练。FICE是具有两个耦合组件的图神经网络(GNN)。首先,在知识图谱的因子图视图上的编码器GNN生成实体和关系嵌入。我们证明,BGP基数是该视图中绑定项周围2跳邻域的局部函数,这为局部消息传递编码器提供了动机。然后,解码器GNN沿着查询的连接拓扑结构组合这些嵌入,以预测对数基数。编码器和解码器联合训练,使嵌入专门用于基数估计。FICE使用邻域采样进行训练,可扩展到具有数百万个三元组的知识图谱,并将嵌入生成与基数解码解耦,以实现低于毫秒的估计延迟。在10个知识图谱上与基于学习和非学习的基线相比,FICE将整体中位数q误差从最佳竞争者的13.54降至5.34,并在尾部行为上优于所有方法。
英文摘要
Query optimization of Basic Graph Patterns (BGP) SPARQL queries over Knowledge Graphs (KG) requires accurate cardinality estimation. Recently published learned estimators outperform statistics- and sampling-based approaches, but share a limitation preventing their adoption in real-world triplestores: they are transductive and require retraining when the underlying graph changes or when applied to new graphs. We present FICE (Fully Inductive Cardinality Estimation), the first learned cardinality estimator for BGP queries over KGs that generalizes to entirely unseen graphs (including unseen relations), without any retraining. FICE is a graph neural network (GNN) with two coupled components. First, an encoder GNN over a factor-graph view of the KG produces entity and relation embeddings. We prove that BGP cardinality is a local function of the 2-hop neighborhood around bound terms in this view, motivating the local message-passing encoder. A decoder GNN then composes these embeddings along the join topology of the query to predict log-cardinality. The encoder and decoder are trained jointly, making the embeddings specialized for cardinality estimation. FICE is trained using neighborhood sampling to scale to KGs with millions of triples, and decouples embedding generation from cardinality decoding to enable estimation latency below a millisecond. Compared to learned and non-learned baselines over 10 KGs, FICE reduces the overall median q-error from 13.54 (for the best competitor) to 5.34 and dominates all approaches in tail behavior.
CommentsExtended version of a paper accepted at ISWC 2026. 34 pages, 8 figures