arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

垂直分区联邦知识图谱的成本表征

Cost Characterization of Vertically Partitioned Federated Knowledge Graphs

Md Saikat Islam Khan Bappy, Oshani Seneviratne

arXiv 2609.13664首次发表:更新:

发表机构

Rensselaer Polytechnic Institute(伦斯勒理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究形式化垂直分区联邦知识图谱的设计空间,比较四种分区策略,发现三个成本指标由图和孤岛数决定,仅跨孤岛路径长度与负载均衡构成权衡,为部署提供指导。

AI 中文摘要

知识图谱日益分布在自治组织之间,这些组织共享实体空间但拥有不相交的关系子集,形成垂直分区。回答多跳查询可能需要结合来自多个数据孤岛的事实,这使得分区策略成为影响通信、索引、负载均衡和查询延迟的关键数据管理决策。然而,与不同分区策略相关的成本仍未得到充分研究。我们将垂直分区形式化为一个设计空间,并比较四种策略:语义域分组、频率平衡分区、共现图割分区和随机分区。我们使用五个指标对其进行评估:通信成本、候选索引大小、跨孤岛路径长度、负载均衡和端到端查询延迟。其中三个指标被证明由图和孤岛数量决定,而非由分区决定,这使设计问题简化为两个冲突的轴:跨孤岛路径长度和负载均衡。在MetaQA和PathQuestion上的实验使用基于TransE嵌入和冻结BERT编码器的固定联邦知识图谱问答架构,跨越三种孤岛配置。通过保持学习模型不变,我们隔离了分区的影响,并表明局部性与均衡性之间的权衡仅在每个孤岛可以容纳多个关系时成立,随着孤岛数量的增加而减弱。该研究为受跨孤岛推理或孤岛负载约束的部署提供了实用指导。

英文摘要

Knowledge graphs are increasingly distributed across autonomous organizations that share an entity space but own disjoint subsets of relations, forming a vertical partition. Answering a multi-hop query may require combining facts from several silos, making the partitioning strategy a key data management decision that affects communication, indexing, load balance, and query latency. However, the costs associated with different partitioning strategies remain insufficiently studied. We formalize vertical partitioning as a design space and compare four strategies: semantic domain grouping, frequency-balanced partitioning, co-occurrence graph-cut partitioning, and random partitioning. We evaluate them using five metrics: communication cost, candidate index size, cross-silo path length, load balance, and end-to-end query latency. Three of the five prove to be determined by the graph and the silo count rather than by the partition, which reduces the design problem to two conflicting axes, cross-silo path length and load balance. Experiments on MetaQA and PathQuestion use a fixed federated knowledge graph question-answering architecture based on TransE embeddings and a frozen BERT encoder across three silo configurations. By keeping the learning model unchanged, we isolate the effect of partitioning and show that the trade-off between locality and balance holds only where each silo can hold several relations, weakening as the number of silos increases. The study provides practical guidance for deployments constrained by cross-silo reasoning or by silo load.

CommentsAccepted at DMKG'26: 2nd International Workshop on Data Management for Knowledge Graphs, co-located with ISWC 2026; to appear in CEUR-WS proceedings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑