发表机构
University of Toronto; UTOA Computing Analytics Co., Ltd.; King Mongkut’s University of Technology Thonburi(多伦多大学; UTOA计算分析有限公司; 国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对法律基准缺乏立场辩护与多条款推理评估的问题,提出GRACE数据集及CLeAR-4B模型,验证了基于法定文本的轻量级法律推理的有效性。
AI 中文摘要
大型语言模型在一系列法律任务中表现出强大的性能,但现有的基准很少评估其提出并捍卫法律立场、在不完整信息下进行推理或综合多个法定条款的能力。这一差距在加拿大法律领域尤为明显,而加拿大法律在法律自然语言处理中仍然代表性不足。我们引入了GRACE(基于加拿大法律的基础对抗推理示例),这是一个包含1,915个基于加拿大联邦立法的问题-推理-答案实例的数据集。GRACE涵盖三种推理模式:对抗性辩护、不确定性和应用推理。我们开发了一个流水线,该流水线对原始法定文本进行分区,生成基于场景的问题和推理,并通过无模型的引文验证和基于LLM的质量审计来过滤示例。作为概念验证,我们对CLeAR-4B(加拿大法律对抗推理)进行了微调,这是一个用于基础法律推理的轻量级模型,并在开放和封闭书籍设置中将其与未修改的Qwen3-4B基础模型进行了评估。当提供相关法案文本时,CLeAR-4B显著提高了与教师输出的一致性以及法定引文行为,而当法规被扣留时,其基础性急剧下降。这些结果表明,GRACE可以支持开发能够更有效地从提供的法定文本中进行推理的轻量级法律模型。
英文摘要
Large language models have shown strong performance across a range of legal tasks, but existing benchmarks rarely evaluate the ability to take and defend a legal position, reason under incomplete information, or synthesize multiple statutory provisions. This gap is particularly pronounced for Canadian law, which remains underrepresented in legal NLP. We introduce GRACE (Grounded Reasoning Adversarial Canadian LEgal examples), a dataset of 1,915 question-reasoning-answer instances grounded in Canadian federal legislation. GRACE covers three reasoning modes: adversarial advocacy, uncertainty, and applied reasoning. We develop a pipeline that partitions raw statutory text, generates scenario-based questions and reasoning, and filters examples through model-free citation verification and LLM-based quality auditing. As a proof of concept, we fine-tune CLeAR-4B (Canadian Legal Adversarial Reasoning), a lightweight model for grounded legal reasoning, and evaluate it against the unmodified Qwen3-4B base model in open- and closed-book settings. CLeAR-4B substantially improves agreement with teacher outputs and statutory citation behavior when the relevant act text is provided, while its grounding degrades sharply when the statute is withheld. These results suggest that GRACE can support the development of lightweight legal models that reason more effectively from supplied statutory text.