RENSA:用于导航共享分布式端点以自动生成联邦SPARQL查询的丰富环境元数据
RENSA: Rich Environment Metadata to Navigate Shared and Distributed Endpoints for Automated Federated SPARQL Query Generation
浏览论文内容
中文总结 AI 辅助
RENSA是利用扩展SBM的联邦SPARQL查询生成框架,可精准选择数据源、消除ASK查询开销,在LargeRDFBench上表现优异,还能为查询变量推断约束。
中文摘要 AI 辅助
随着知识图谱技术的普及,知识图谱数据库的数量大幅增长,知识图谱可通过联邦SPARQL查询实现分布式数据的动态集成。然而,由于去中心化数据集缺乏详细的结构知识,在联邦环境中构建高效查询颇具挑战。尽管VoID等标准提供了基础元数据,但往往无法捕捉优化所需的复杂链接和权威分布,导致当前引擎常依赖运行时ASK查询进行源选择,增加了通信开销。我们提出RENSA,这是一种联邦SPARQL查询生成框架,利用SPARQL构建器元数据(SBM)的扩展版本,通过整合类和权威信息,将主语和对象使用映射到特定谓词,使RENSA无需运行时通信即可实现精准的源选择和查询变量的语义约束推理。生成的配置文件在大多数情况下仅占原始数据集三元组的不到1%,确保了存储效率。在LargeRDFBench基准(13个数据集,超过10亿个三元组,32个查询)上的评估显示,RENSA的源选择结果可与最先进方法媲美,同时消除了ASK查询开销;此外,我们证明RENSA可为查询变量推断类和权威约束,甚至能跨异构端点识别数据源,这些配置文件还为半自动查询生成提供了人类可读的结构洞察。
英文摘要
The number of knowledge graph databases has increased significantly with the proliferation of knowledge graph technologies. Knowledge graphs enable the dynamic integration of distributed data through federated SPARQL queries. However, constructing efficient queries in a federated environment is challenging due to the lack of detailed structural knowledge across decentralized datasets. While standards like VoID provide basic metadata, they often fail to capture the complex interlinks and authority distributions necessary for optimization. Consequently, current engines frequently rely on runtime ASK queries for source selection, increasing communication overhead. We propose RENSA, a federated SPARQL query generation framework that leverages an extension of SPARQL Builder Metadata (SBM). By integrating class and authority information, mapping subject and object usage to specific predicates, RENSA enables precise source selection and semantic constraint inference for query variables without runtime communication. The generated profiles represent less than 1\% of the original dataset triples in most cases, ensuring storage efficiency. Evaluation on the LargeRDFBench benchmark (13 datasets with >1B triples, 32 queries) shows that RENSA achieves source selection results comparable to state-of-the-art methods while eliminating ASK query overhead. Furthermore, we demonstrate that RENSA infers class and authority constraints for query variables, enabling the identification of data sources even across heterogeneous endpoints. These profiles additionally offer human-readable structural insights for semi-automated query generation.
发表机构
- Graduate University for Advanced Studies, SOKENDAI(综合研究大学院大学(SOKENDAI))
- National Institute of Informatics(信息学研究所)
- Database Center for Life Science(生命科学数据库中心)
机构由 AI 辅助整理,请以论文原文为准。