arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11701cs.DB

借助地图行进:利用RDF形状减少链接遍历查询的搜索空间

Traveling with a Map: Reducing the Search Space of Link Traversal Queries Using RDF Shapes

  • Ghent University – imec(根特大学–imec)

机构由 AI 辅助整理,请以论文原文为准。

Bryan-Elliott Tam, Joanna Van Herwegen, Pieter Colpaert, Ruben Verborgh, Ruben Taelman

AI总结:

针对去中心化网络中LTQP执行慢、成本高的问题,提出基于RDF形状的剪枝方法,可显著优化数据模型选择性查询的处理效率。

AI中文摘要:

网络信息的集中化引发了法律和伦理方面的担忧,尤其在社交、医疗和教育应用领域。去中心化架构通过将数据保留在其来源附近提供了一种有前景的替代方案,但高效的查询处理仍然是一项重大挑战。链接遍历查询处理(LTQP)支持跨去中心化网络进行查询,但由于涉及大量HTTP请求,往往存在执行时间长、数据传输成本高的问题。许多查询对于分布在网络中的数据模型对象具有高度选择性,例如在社交媒体应用中,用户存储异构数据,查询可能仅针对用户的帖子和评论,而忽略其他信息,我们将此类查询称为数据模型选择性查询。我们提出一种基于形状的剪枝方法,该方法依赖于形状索引和查询-形状包含算法来减少搜索空间,从而减少HTTP请求的数量。我们将此方法形式化为LTQP的链接剪枝机制,并在SolidBench基准测试的社交媒体查询上针对多个指标进行评估。结果表明,基于形状的剪枝可显著提升数据模型选择性查询的执行时间、首结果到达时间、效率和网络使用量,而对非选择性数据模型查询的影响可忽略不计。这些增益仅导致每个形状索引实例的三元组数量略有增加。我们的方法具有鲁棒性,即使部分数据提供者未提供形状索引,仍能保留其优势。本研究表明,基于形状的元数据可显著优化去中心化知识图中针对重要查询类别的LTQP,数据提供者通过公开此类元数据不仅能提升数据质量和互操作性,还能提高基于遍历的查询处理的效率。

英文摘要:

The centralization of web information raises legal and ethical concerns, particularly in social, healthcare, and education applications. Decentralized architectures offer a promising alternative by keeping data closer to its source, yet efficient query processing remains a significant challenge. Link Traversal Query Processing (LTQP) enables querying across decentralized networks but often suffers from long execution times and high data transfer costs due to the large number of HTTP requests involved. Many queries are highly selective with respect to the data model objects distributed across the network. For example, in a social media application where users store heterogeneous data, a query may target only users' posts and comments, ignoring their other information. We refer to such queries as data-model selective. We propose a shape-based pruning approach that relies on shape indexes and a query-shape subsumption algorithm to reduce the search space and thus the number of HTTP requests. We formalize this approach as a link pruning mechanism for LTQP and evaluate it on social media queries from the SolidBench benchmark across multiple metrics. Our results show that shape-based pruning substantially improves query execution time, first-result arrival time, diefficiency, and network usage for data-model selective queries, while having a negligible impact on non-selective data-model queries. These gains cost only a minor increase in triples per shape-index instance. Our approach is also resilient, retaining its benefits even when some data providers do not supply shape indexes. This work demonstrates that shape-based metadata can significantly optimize LTQP in decentralized knowledge graphs for an important class of queries. By exposing such metadata, data providers not only enhance data quality and interoperability but also improve the efficiency of traversal-based query processing.

↑