arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

法律大语言模型幻觉应被评估为法律授权(Legal Warrant)的失败

Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant

Maksym Taranukhin, Vered Shwartz

arXiv 2609.17546首次发表:更新:

AI 中文总结

本文主张法律大模型幻觉应评估为法律授权失败,提出主张-权威授权关系及相应评估指标,并给出基准设计、评分方法和研究议程,以检验系统主张是否获得法律许可。

AI 中文摘要

在这篇立场论文中,我们认为法律大语言模型(LLM)的幻觉应被评估为法律授权(legal warrant)的失败,而非事实不准确或引用失败。我们将主张-权威授权(claim-authority warrant)定义为一种上下文敏感的关系,即一个具有重要后果的法律主张与权威之间的关系,该权威存在、适用于相关司法管辖区、在分析日期具有现行效力、具有系统所表示的法律地位,并支持所断言的主张。有授权的法律生成(warranted legal generation)是更广泛的系统行为,它根据这种关系进行回答、缩小范围、提问、警告、纠正错误前提或弃权(不执行)。可证伪的预测是,授权指标(warrant metrics)能够揭示实质性失败,而这些失败是回答准确性、引用存在性、通用归因、LegalHalBench式法规相关性以及CitaLaw式句子-引用对齐可能遗漏的。我们通过一个并排比较项目和一项在公开规则测试上可复现的小型试点来强化这一主张。随后,我们详细说明了基准记录、主张边界、支持标签、混合响应策略评分、风险权重、注释可靠性报告以及特定司法管辖区的权威本体。其结果是提出了一项具体的研究议程,用于根据法律AI系统的具有重要后果的主张是否得到法律许可来评估这些系统。

英文摘要

In this position paper, we argue that legal LLMs' hallucinations should be evaluated as a failure of legal warrant rather than as factual inaccuracy or citation failure. We define claim-authority warrant as the context-sensitive relation between a consequential legal claim and authority that exists, applies to the relevant jurisdiction, is current for the date of analysis, has the legal status represented by the system, and supports the proposition asserted. Warranted legal generation is the broader system behavior that answers, narrows, asks, warns, corrects a false premise, or abstains according to that relation. The falsifiable prediction is that warrant metrics reveal material failures that answer accuracy, citation existence, generic attribution, LegalHalBench-style statute relevance, and CitaLaw-style sentence-citation alignment can miss. We sharpen this claim with a side-by-side comparison item and a small, reproducible pilot over public-rule tests. We then specify benchmark records, claim boundaries, support labels, mixed response-policy scoring, risk weights, annotation reliability reporting, and jurisdiction-specific authority ontologies. The result is a concrete research agenda for evaluating legal AI systems by whether their consequential claims are licensed by law.

CommentsAccepted at AI4Law@ICML2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑