arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LAURA:法律合同中可解释歧义条款识别的知识蒸馏

LAURA: Knowledge Distillation for Interpretable Ambiguous Clause Identification in Legal Contracts

Amrita Singh, Aditya Joshi, Jiaojiao Jiang, Hye-young Paik

arXiv 2609.36707首次发表:更新:

发表机构

University of New South Wales (UNSW)(新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LAURA通过知识蒸馏和IRAC-Unlearning提示技术,将教师大语言模型知识迁移至小型开放权重模型,实现法律合同歧义条款的可解释识别,兼顾识别性能与可解释性。

AI 中文摘要

法律合同中的歧义会使企业面临财务和法律风险。有些歧义允许灵活解释而不会引发争议,而另一些则会导致重大的法律冲突。这使得仅识别歧义是不够的,可解释的理据分析至关重要。我们提出了LAURA,一个用于可解释歧义条款识别的后训练框架。LAURA利用知识蒸馏和一种IRAC-Unlearning提示技术,将知识从教师大语言模型迁移到一个开放权重的学生模型(参数不超过10亿),随后使用结合分类和理据生成损失的联合目标对该学生模型进行训练。该框架支持法律和非法律利益相关者就哪些歧义需要进一步关注做出明智决策。在7个基线和7个开放权重模型上的大量实验表明,使用Flan-T5(250M)的LAURA在所有可解释基线中达到了最先进的可解释性,同时匹配了性能最佳的不透明基线的识别性能。

英文摘要

Legal contracts contain ambiguities that expose enterprises to financial and legal risks. Some ambiguities allow flexible interpretation without triggering disputes, while others lead to significant legal conflicts. This makes identification alone insufficient, and interpretable rationale analysis essential. We propose LAURA, a post-training framework for interpretable ambiguous clause identification. LAURA leverages knowledge distillation with an IRAC-Unlearning prompting technique to transfer knowledge from a teacher LLM to an open-weight student model (<=1B parameters), which is then trained using a joint objective combining classification and rationale generation losses. The framework supports both legal and non-legal stakeholders in making informed decisions about which ambiguities require further attention. Extensive experiments across 7 baselines and 7 open-weight models demonstrate that LAURA with Flan-T5 (250M) delivers state-of-the-art interpretability over all interpretable baselines while matching the identification performance of the best-performing opaque baseline.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑