发表机构
Bloomberg(彭博公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对现有大语言模型法律用例评估忽视的“指向真实案件却不支持主张”的引用问题,通过扰动真实法律引用验证14种模型配置,发现模型易混淆主题识别与页码级支持验证能力,规模与推理能力仅能部分缩小错精准页码扰动的检测差距。
AI 中文摘要
2023年,纽约一名法官在Mata诉Avianca案中对两名律师进行了制裁,原因是他们提交的辩护状中包含由ChatGPT生成的虚构引用。这类错误大多可以通过数据库查找被发现;更棘手的问题是检测那些指向真实案件但不支持所提出主张的引用——这种失败模式在现有针对法律用例的大语言模型(LLM)评估中基本被忽视。在本文中,我们通过对从两个法律语料库中获取的真实法律引用进行可控扰动来研究主张级引用支持验证,扰动方式要么是替换被引用的案件,要么仅更改同一案件内的精准页码。我们在生成的示例上评估了14种模型配置。模型能捕获93%至100%的错案扰动;对于错精准页码扰动,模型在法院意见上仅能捕获37%至61%,在法律辩护状上则为52%至83%。当模型未能捕获错精准页码扰动时,它们是基于主题重叠而非页码级支持来接受该引用。规模和扩展推理缩小了差距但并未消除差距:高推理能力的GPT-5.4在法院意见上仍会遗漏40%的精准页码不匹配,在辩护状上则遗漏18%。提示模型验证被引用页面的支持情况可提高召回率,但也会增加误报率。识别正确的法律主题和验证对所引用主张的支持是两种不同的能力,而当前模型将二者混淆。
英文摘要
In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations generated by ChatGPT. Such failures are largely caught by database lookups; the harder problem is detecting citations that point to real cases but do not support the propositions for which they are offered -- a failure mode that existing evaluations of LLMs for legal use cases largely overlook. In this paper, we study proposition-level citation support verification through controlled perturbations of real legal citations obtained from two legal corpora, either replacing the cited case or changing only the pinpoint page within the same case. We evaluate fourteen model configurations on the resulting examples. Models catch 93-100% of wrong-case corruptions. They catch only 37-61% of wrong-pinpoint corruptions on court opinions and 52-83% on legal briefs. When models fail to catch wrong-pinpoint corruptions, they accept the citation based on topical overlap rather than page-level support. Scale and extended reasoning narrow the gap but do not close it: GPT-5.4 with high reasoning effort still misses 40% of pinpoint mismatches on court opinions and 18% on briefs. Prompting the model to verify support at the cited page improves recall, but it also raises the false positive rate. Recognizing the right legal topic and verifying support for the cited proposition are distinct capabilities, and current models conflate them.
CommentsAccepted to the 1st Workshop on AI for Law at ICML 2026