arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Build2SPARQL:面向建筑知识图谱查询的大规模文本到SPARQL基准数据集

Build2SPARQL: A Large-Scale Text-to-SPARQL Benchmark Dataset for Building Knowledge Graph Querying

Wooyoung Jung

arXiv 2610.00224首次发表:更新:

AI 中文总结

提出Build2SPARQL,一个由知识图谱接地流程生成的大规模建筑知识图谱文本到SPARQL基准,包含6,136个查询和30,680个问题,人工验证高保真度,检索增强评估显著提升模型准确率。

AI 中文摘要

建筑自动化系统越来越多地使用诸如Brick和ASHRAE 223P等本体表示为语义知识图谱(KGs),为人工智能应用创造了机器可读的基础。一个有前景的应用是将自然语言问题转换为SPARQL(文本到SPARQL),这将使建筑运营商能够通过语言代理查询这些图谱,但进展受到大型自然语言/SPARQL基准稀缺的限制。本文提出了Build2SPARQL,一个由知识图谱接地流程生成的大规模建筑知识图谱基准:SPARQL查询完全由图遍历代码生成和验证,而大型语言模型仅生成自然语言问题,从而保持查询正确性独立于模型行为。该流程挖掘六种查询模式家族——线性链、分支、UNION、聚合、OPTIONAL和属性过滤——并在五种词汇语域中表述每个查询。应用于201个建筑知识图谱(180个Brick,21个ASHRAE 223P),产生了6,136个可执行的SPARQL查询和30,680个问题。对300个问题进行的双评分者人工验证发现,语义保真度为98.8%,自然度为98.8%,操作合理性为84.0%。跨三个开放权重语言模型的检索增强评估将精确匹配准确率从0.2-20%(零样本)提高到56-65%(三样本检索)。

英文摘要

Building automation systems are increasingly represented as semantic knowledge graphs (KGs) using ontologies such as Brick and ASHRAE 223P, creating a machine-readable substrate for artificial-intelligence applications. One promising application is translating natural-language questions into SPARQL (text-to-SPARQL), which would let building operators query these graphs through language agents, but progress is limited by the scarcity of large natural-language/SPARQL benchmarks. This paper presents Build2SPARQL, a large-scale benchmark for building KGs generated by a KG-grounded pipeline: SPARQL queries are produced and validated entirely by graph-traversal code, while large language models generate only the natural-language questions, keeping query correctness independent of model behavior. The pipeline mines six query-pattern families -- linear chains, branching, UNION, aggregation, OPTIONAL, and attribute-filtered -- and phrases each query across five vocabulary registers. Applied to 201 building KGs (180 Brick, 21 ASHRAE 223P), it yields 6,136 executable SPARQL queries and 30,680 questions. A two-rater human validation of 300 questions found 98.8% semantic fidelity, 98.8% naturalness, and 84.0% operational plausibility. A retrieval-augmented evaluation across three open-weight language models raised exact-match accuracy from 0.2-20% (zero-shot) to 56-65% (three-shot retrieved).

Comments26 pages, 2 figures, 16 tables. Data paper. Dataset openly available at https://zenodo.org/records/20451474. Under review at the ASCE Journal of Computing in Civil Engineering

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑