arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CQ4OE:一个用于评估基于能力问题的大语言模型辅助本体生成的基准

CQ4OE: A benchmark for assessing LLM-assisted ontology generation from competency questions

Jiayi Li, Ziyuan Wang, Daniel Garijo, María Poveda-Villalón

arXiv 2609.26029首次发表:更新:

发表机构

Ontology Engineering Group, Universidad Politécnica de Madrid(马德里理工大学本体工程组)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CQ4OE是一个新基准,用于系统评估大语言模型从能力问题生成本体的性能,提供带来源的黄金本体和两个互补任务,实验表明LLM在术语恢复上优于完整本体构建。

AI 中文摘要

从能力问题(CQs)生成本体是本体工程中一个核心但劳动密集的阶段。尽管大语言模型(LLMs)提供了有前景的自动化能力,但当前的评估仍然零散。任务表述异构,黄金标准往往缺乏细粒度的CQ来源,指标将词汇重叠与结构和逻辑充分性混为一谈,且参考本体并非总是明确围绕评估CQs设计。在此,我们通过CQ4OE解决这些局限,这是一个用于系统且可重复地评估基于LLM的从CQs生成本体的基准。对于基准中的每个本体,我们构建了一个CQ驱动的黄金OWL本体,具有明确的来源,将每个CQ与回答该问题所需的类、属性和公理联系起来。从该资源中,我们定义了两个互补的评估任务。CQ2Term支持对99个CQs的CQ特定类和属性预测进行术语级评估,CQ2Onto支持对118个CQs的本体级评估,包括层次结构、属性建模和公理级结构。我们通过使用九种LLM在零样本、迭代和多智能体生成策略下的实验展示了CQ4OE,表明LLM恢复显式词汇术语比创建本体更可靠,尤其是在属性建模、层次结构构建和公理生成方面。

英文摘要

Ontology generation from Competency Questions (CQs) is a central yet labor-intensive phase of Ontology Engineering. While large language models (LLMs) offer promising automation capabilities, current evaluations remain fragmented. Task formulations are heterogeneous, gold standards often lack fine-grained CQ provenance, metrics conflate lexical overlap with structural and logical adequacy, and reference ontologies are not always explicitly designed around the evaluation CQs. Here, we address these limitations with CQ4OE, a benchmark for the systematic and reproducible evaluation of LLM-based ontology generation from CQs. For each ontology in the benchmark, we build a CQ-driven gold OWL ontology with explicit provenance linking each CQ to the classes, properties, and axioms required to answer it. From this resource, we define two complementary evaluation tasks. CQ2Term supports term-level evaluation of CQ-specific class and property prediction over 99 CQs, and CQ2Onto supports ontology-level evaluation over 118 CQs, including hierarchy, property modeling, and axiom-level structure. We demonstrate CQ4OE with experiments using nine LLMs under zero-shot, iterative, and multi-agent generation strategies, showing that LLMs recover explicit vocabulary terms more reliably than creating ontologies, particularly in property modeling, hierarchy construction, and axiom generation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑