arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2608.07530cs.AIcs.CLcs.DB

NL2SHACL-Bench:面向自然语言到SHACL翻译的基准套件

NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation

  • Technical University of Munich(慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuchen Zhou, Niels Bobet, Maribel Acosta

AI总结:

本文提出NL2SHACL-Bench基准套件,评估了四个最先进大语言模型的NL2SHACL翻译能力,发现当前模型能生成语法有效SHACL,但复杂逻辑与结构模式的语义等价约束生成仍有不足,该基准可衡量领域进展。

AI中文摘要:

SHACL是验证RDF知识图谱(KGs)一致性的核心技术。然而,编写SHACL形状需要领域专家大多缺乏的技术专业知识。将自然语言需求翻译为SHACL(NL2SHACL)可降低这一门槛,但目前尚无针对NL2SHACL的专用基准,且评估生成的形状需采用超出字符串比较的方法,因为语义等价的形状可能在序列化和结构上存在差异。为应对这些挑战,本文提出NL2SHACL-Bench,这是一个面向自然语言到SHACL翻译的基准套件。利用NL2SHACL-Bench,我们评估了四个针对该任务的最先进大语言模型(LLMs)。结果显示,当前LLMs具备生成语法有效SHACL的较强能力,但在针对复杂逻辑和结构模式生成语义等价约束方面仍存在困难。这表明NL2SHACL-Bench为衡量NL2SHACL领域的进展提供了有意义的基础。

英文摘要:

SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical expertise that most domain experts lack. Translating natural language requirements into SHACL (NL2SHACL) would lower this barrier. However, there is no dedicated benchmark for NL2SHACL, and evaluating generated shapes requires methods beyond string comparison, as semantically equivalent shapes can differ in serialisation and structure. To tackle these challenges, we present NL2SHACL-Bench, a benchmark suite for natural language to SHACL translation. Using NL2SHACL-Bench, we evaluate four state-of-the-art large language models (LLMs) for this task. Our results show that current LLMs are highly capable of generating syntactically valid SHACL, but still struggle to produce semantically equivalent constraints for complex logical and structural patterns. This indicates that NL2SHACL-Bench provides a meaningful basis for measuring advances in the NL2SHACL state of the art.

补充信息

↑