arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12326cs.CL

论法律本体学习中的语义保留度量

On Measuring Semantic Preservation in Legal Ontology Learning

  • Warsaw University of Technology(华沙理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Albert Sadowski, Jarosław A. Chudziak

AI总结:

本文针对法律本体学习的语义保留评估问题,提出通过对比LLM在源文档与转换后表示上的任务性能量化语义损失,经多模型多方法实验发现语义损失随模型-方法配对显著变化,为法律知识系统配置选择提供指导。

AI中文摘要:

本体学习将非结构化文本转换为结构化表示以实现自动推理,但结构化信息的过程可能导致信息丢失,而当前评估方法仅关注结构正确性,无法检测此类丢失,也无法衡量转换后意义是否保留。本文提出一种解决该问题的评估方法:对比源文档与转换后表示上的大语言模型(LLM)任务性能,两者的差异可量化语义损失。我们将该方法应用于法律并购协议分析领域,该领域因语言复杂、语义要求精确而被选中,对比了直接应用LLM与三种本体学习方法在六个语言模型上的表现。结果显示存在系统性语义损失,且损失程度随推理复杂度、模型与方法的交互作用存在显著差异。本文的贡献为:(1)提出用于度量本体学习中语义保留的评估框架;(2)提供实证证据表明语义损失随模型-方法配对存在显著变化,为法律知识系统中最优配置的选择提供指导。

英文摘要:

Ontology learning transforms unstructured text into structured representations for automated reasoning. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on structural correctness while failing to measure whether meaning survives transformation. We propose an evaluation methodology that addresses this: comparing LLM task performance on source documents against performance on transformed representations, with the difference quantifying semantic loss. We demonstrate this approach on legal merger agreement analysis, a domain chosen for its complex language and precise semantic requirements, comparing direct LLM application against three ontology learning methods across six language models. The results reveal systematic semantic loss with significant variation based on reasoning complexity and model-method interactions. Our contributions are: (1) an evaluation framework for measuring semantic preservation in ontology learning, and (2) empirical evidence that semantic loss varies dramatically with model-method pairing, providing guidance for selecting optimal configurations in legal knowledge systems.

补充信息

↑