发表机构
Thomson Reuters Foundational Research; Imperial College London(汤森路透基础研究部; 帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对合同 scrubbing 任务,研究人员推出首个基准测试 ContractScrub,发现前沿 LLMs 在该任务上表现不佳,凸显了领域特定基准的重要性。
AI 中文摘要
法律工作高度依赖大量文本处理,被认为是最适合使用大语言模型(LLMs)的领域之一。合同“ scrubbing”即对交易协议进行最终审核以发现错误和不一致之处,是一项特别适合自动化的任务,因为它是日常性的、需要对长文档进行细致关注的艰苦工作。此外,合同 scrubbing 似乎与前沿 LLMs 所具备的长上下文推理、一致性检查和命名实体识别(NER)等通用能力天然契合。尽管合同 scrubbing 具有经济价值和自动化潜力,但目前尚未对执行合同 scrubbing 的 LLMs 进行正式评估。我们推出 ContractScrub,这是首个用于评估合同 scrubbing 能力的基准测试,包含由经验丰富的律师手工制作的合同,涵盖误用定义术语、引用错误、语言不一致等多种错误类别。前沿模型的表现差得出乎意料,尽管在看似相关的通用基准测试中表现强劲,但仅有一个模型达到了 0.75 的宏平均召回率,这表明了当前模型的实际局限性,以及针对特定领域的精准基准测试对于衡量实际影响的重要性。
英文摘要
Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final review of transactional agreements for errors and inconsistencies, is a particularly suitable task for automation, because it is routine, painstaking work requiring detailed attention to long documents. Scrubbing also seems to align naturally with the general capabilities expected of frontier LLMs around long-context reasoning, consistency checking, and named entity recognition (NER). Despite the economic value and potential for automation, no formal evaluations of LLMs performing contract scrubbing have been conducted. We introduce ContractScrub, the first benchmark designed to evaluate contract scrubbing capabilities, comprising contracts hand-crafted by experienced lawyers over diverse error categories such as misuse of defined terms, incorrect references, and inconsistent language. Frontier models perform surprisingly poorly with only one model reaching 0.75 macro average recall despite strong performance on seemingly related general benchmarks, demonstrating the practical limits of current models and the importance of narrowly targeted, domain-specific benchmarks for measuring real-world impact.
Comments10 pages, ICML AI4Law Workshop