GrOIL:基于图的领域本体归纳与受限大语言模型调解
GrOIL: Graph-Grounded Domain Ontology Induction with Constrained LLM Mediation
浏览论文内容
中文总结 AI 辅助
该研究提出GrOIL流程,以受限LLM调解的七阶段图基方法构建可审计OWL TBox,在人寿保险领域的基准测试中,其CQ覆盖率等指标优于LLM基线,可生成稳定领域表示。
中文摘要 AI 辅助
从领域文档构建形式化本体,需同时实现语料库接地、词汇一致性、公理级表达能力及端到端可追溯性,现有自动系统均无法同时满足这四个要求。本文提出一个七阶段基于图的流程,可将领域文档转换为完整、可审计的Web本体语言(OWL)术语盒(TBox),且无任何无约束生成步骤。文档首先被编码为统一话语超图(UDH),以捕获实体参与关系与话语依赖;后续阶段将此图证据转换为类层次结构、带类型的对象属性与数据属性,以及约束公理,其中大语言模型(LLM)仅用于狭窄的、基于图的调解任务。配对的断言盒(ABox)填充流程将命名个体接地到归纳出的TBox,支持基于SPARQL的功能评估。每个生成的术语都带有从原始源段落到各流程阶段的完整决策链,使TBox可直接审计并适合针对性人工优化。在人寿保险领域,使用两个已建立的基准及一个包含10种产品类型的100份合同的新语料库进行评估,我们的流程在所有评估维度均取得优异结果:在能力问题(CQ)覆盖率上,其表现优于直接LLM基线与多智能体LLM基线(一份定期寿险合同上为0.85,而两种基线分别为0.63和0.62;另一份合同上为0.77,两种基线分别为0.40和0.44);同时实现了与人工构建的参考本体相当的高关键词覆盖率,以及在结构化缺口与重叠推理上的优异性能,且无需任何手动TBox工程。本体增长分析提供了与大规模词汇饱和一致的证据,表明该流程可从大型文档语料库生成稳定、可复用的领域表示。
英文摘要
Constructing formal ontologies from domain documents requires simultaneously enforcing corpus grounding, vocabulary consistency, axiom-level expressivity, and end-to-end provenance, a combination no existing automatic system delivers. We present a seven-stage graph-grounded pipeline that converts domain documents into a complete, auditable Web Ontology Language (OWL) Terminological Box (TBox) without any unconstrained generation step. Documents are first encoded as Unified Discourse-Hypergraphs (UDH) capturing entity participation and discourse dependencies; subsequent stages transform this graph evidence into a class hierarchy, typed object and datatype properties, and restriction axioms, with Large Language Model (LLM) usage restricted to narrow, graph-grounded mediation tasks. A paired Assertional Box (ABox) population procedure grounds named individuals in the induced TBox, enabling SPARQL-based functional evaluation. Every emitted term carries a full decision chain from raw source passages through each pipeline stage, making the TBox directly auditable and suitable for targeted human refinement. Evaluated on the life insurance domain using two established benchmarks and a new 100-contract corpus spanning ten product types, our pipeline achieves strong results across all evaluation dimensions, outperforming direct and multi-agent LLM baselines on competency-question (CQ) coverage (0.85 vs. 0.63 and 0.62 on one term-life contract; 0.77 vs. 0.40 and 0.44 on another contract), while also attaining high keyphrase coverage comparable to a manually-constructed reference ontology and strong performance on structured gap-and-overlap reasoning, all without any manual TBox engineering. Ontology growth analysis provides evidence consistent with vocabulary saturation at scale, demonstrating that the pipeline produces stable, reusable domain representations from large document corpora.