超越准确性与表面流畅性:面向法律条款生成的LLM风险敏感评估
Beyond Accuracy and Surface Fluency: Risk-Sensitive Evaluation of LLMs for Legal Clause Generation
浏览论文内容
中文总结 AI 辅助
针对LLM生成法律条款时传统评估忽视法律风险的问题,提出结合CLAUSE与LENS-CRAFT框架并采用最大严重性原则的评估方法,实现对条款质量的风险敏感评估。
中文摘要 AI 辅助
大型语言模型(LLMs)越来越多地被用于起草合同语言,然而传统的准确性或基于偏好的评估方式与法律起草工作并不匹配。一个条款可能语言流畅、风格精良,却仍然遗漏了关键的例外条款,以不可执行的方式分配风险,假设了不适用的司法管辖区,或使一方暴露于监管责任之下。本文提出了一种用于评估LLM生成的合同条款的实证研究设计和框架。该研究评估了四个模型——Claude Haiku 4.5、Gemini 2.5 Flash Lite、GPT 5.4 Nano和Qwen 3.5 Flash,覆盖22个合同条款类别和34种基于法律动机的失败模式。我们结合了两种评估框架:CLAUSE,它根据法律功能和失败目标对提示进行分类;以及LENS-CRAFT,它在九个法律质量维度上对输出进行评分。该研究并非对各维度分数取平均值,而是应用了最大严重性原则,使得单个具有法律决定性的缺陷仍然可见。本文提供了评估协议、分类体系、分析计划以及用于报告实证结果的结果结构。我们认为,法律AI评估应超越总体准确性,转向针对具体条款、以失败模式为导向且具有风险敏感性的评估。
英文摘要
Large language models (LLMs) are increasingly used to draft contractual language, yet conventional accuracy or preference-based evaluations are poorly matched to legal drafting. A clause may be fluent and stylistically polished while still omitting an essential carve-out, allocating risk in an unenforceable way, assuming an inapplicable jurisdiction, or exposing a party to regulatory liability. This paper presents a empirical study design and framework for evaluating LLM-generated contract clauses. The study evaluates four models - Claude Haiku 4.5, Gemini 2.5 Flash Lite, GPT 5.4 Nano, and Qwen 3.5 Flash, across 22 contract clause categories and 34 legally-motivated failure modes. We combine two evaluation frameworks: CLAUSE, which classifies prompts by legal function and failure target, and LENS-CRAFT, which scores outputs across nine legal-quality dimensions. Instead of averaging dimension scores, the study applies a Max Severity Principle so that a single legally decisive defect remains visible. The paper provides the evaluation protocol, taxonomy, analysis plan, and a results structure for reporting empirical findings. We argue that legal AI evaluation should move beyond aggregate accuracy toward clause-specific, failure-mode-driven, and risk-sensitive assessment.
发表机构
- AI Tech Ethics(AI科技伦理)
机构由 AI 辅助整理,请以论文原文为准。