发表机构
University at Albany, SUNY(纽约州立大学奥尔巴尼分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型语言模型在法律场景中易产生无依据主张的问题,研究构建了PARCEL基准数据集,将任务转化为三向自然语言推理,评估发现模型存在误判缺陷,该基准可用于测试法律RAG系统的主张依据性。
AI 中文摘要
大型语言模型正越来越多地用于法律研究和文书起草,但它们仍可能产生听起来有说服力却未被引用来源支持的主张。我们推出PARCEL,一个用于核查法律主张是否得到基础权威依据支持的基准。利用纽约州上诉法院近期的裁决,我们构建了包含3396个括号式主张的数据集,标注为“支持”“反驳”或“未发现”。我们将该任务转化为三向自然语言推理问题,并在零样本设置下评估多个最先进的大型语言模型。尽管最强模型的准确率可达0.97,但结果也显示出一个重要缺陷:即便提供完整的裁决文本,模型仍会错误地将未被支持的主张标记为已支持。在所有模型中,缺失支持比直接反驳更难检测,而虚构但看似合理的引用会导致性能出现最大幅度的下降。总体而言,PARCEL为测试法律检索增强生成(RAG)系统中的主张级 groundedness(依据性)提供了实用基准。
英文摘要
Large language models are increasingly used in legal research and drafting, but they can still produce claims that sound convincing without being supported by the cited source. We introduce PARCEL, a benchmark for checking whether a legal claim is supported by the underlying authority. Using recent New York State Court of Appeals decisions, we build a dataset of 3,396 parenthetical-style claims labeled as Supported, Refuted, or Not Found. We cast this task as a three-way natural language inference problem and evaluate several state-of-the-art LLMs in a zero-shot setting. Although the strongest models reach up to 0.97 accuracy, the results also show an important weakness: models still incorrectly mark unsupported claims as supported, even when the full opinion text is provided. Across models, missing support is harder to detect than direct contradiction, and fabricated but plausible citations cause the largest drop in performance. Overall, PARCEL provides a practical benchmark for testing claim-level groundedness in legal RAG systems.
Comments9 pages, 2 figures, 7 tables. Published in the 21st International Conference on Artificial Intelligence and Law (ICAIL 2026), Singapore. Dataset: https://github.com/mmikaildemir/PARCEL
Journal refIn Proceedings of the 21st International Conference on Artificial Intelligence and Law (ICAIL 2026), June 08-12, 2026, Singapore. ACM, New York, NY, USA