发表机构
Leiden University(莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文评估开源生成式大语言模型在法律自然语言推理任务上的表现,发现Gemma-4 26B达到81.2%准确率,证明零样本开源模型在无监督数据时是可行替代方案。
AI 中文摘要
在本文中,我们评估了开源生成式大语言模型在法律自然语言推理(NLI)任务上的表现。法律检查过程发生在特定领域,且经常涉及机密数据,这要求使用不需要标注训练数据的本地模型。我们在ContractNLI基准和两个NLI4Wills数据集上评估了我们的模型。我们成功复现了该任务的基线(Span NLI BERT),并在同一任务上评估了多个开源大语言模型。我们分析了模型的无效率,以及它们在不同温度设置和领域间的稳定性。在生成式模型中,Gemma-4 26B表现最佳,达到了81.2%的准确率,甚至在一个指标上超过了监督模型。在准确率方面,零样本方法无法超越监督模型。Qwen-3.6 35B在ContractNLI和法律遗嘱领域的额外数据集上均表现良好。我们的研究结果表明,在没有监督数据可用的情况下,零样本、开源、生成式大语言模型是实际法律NLI任务中可行的替代方案。我们的代码可在以下网址获取:https URL。
英文摘要
In this paper, we evaluate open-source generative LLMs on legal Natural Language Inference (NLI). Legal inspectorial processes take place in specific domains and often deal with confidential data. This creates a need for working with local models that do not require labeled training data. We evaluate our models on the ContractNLI benchmark and two NLI4Wills datasets. We successfully reproduce the baseline for the task (Span NLI BERT) and we evaluate multiple open-source LLMs on the same task. We analyze the invalid rate of the models, and their stability across temperature settings and domains. Among the generative models, Gemma-4 26B performs the best, reaching an accuracy of 81.2%, even outperforming the supervised model on one metric. On accuracy, it is not possible to beat the supervised model with zero-shot approaches. Qwen-3.6 35B performs well on both ContractNLI and additional datasets in the legal wills domain. Our findings indicate that zero-shot, open-source, generative LLMs are a viable alternative for real-world legal NLI when no supervised data is available. Our code is available at https://github.com/fbaratov/contractnli-llms.