arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01292cs.CL

CrossLex:面向大语言模型跨司法辖区法律推理的、基于法律来源的基准测试集

CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models

Xiaocui Yang, Xican Tan, Shoujie Chen, Shihan Xiao, Keke Tong, Xinyu Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出基于中、加州、德三国法律来源的CrossLex基准测试集,定义三项任务并提出联合评估指标,发现大语言模型在跨司法辖区法律推理上存在不足,旨在推动相关研究。

中文摘要 AI 辅助

法律推理本质上依赖于司法辖区:相同的事实在不同法律体系中可能适用不同的法律规则并产生不同结论。然而现有基准测试集很少评估大语言模型(LLM)能否识别这类司法辖区特有的差异,尤其是当相同事实模式导致法律结论分歧时。本文提出CrossLex,这是一个基于相同事实、基于法律来源的基准测试集,用于评估大语言模型在三个司法辖区(中国、美国加利福尼亚州、德国)的跨司法辖区法律推理能力。CrossLex构建于权威法律来源之上,涵盖合同、消费者、刑事、家庭和劳动法共55个法律问题,构建了与司法辖区对齐的问题、答案及支持性引用。CrossLex共包含6149个实例,分为385个事实组,所有法律问题、答案及引用的权威资料均经法律专家审查。为区分基础法律知识与跨司法辖区推理,CrossLex定义了三个互补任务:单一司法辖区推理(T1)、联合跨司法辖区比较(T2)、细粒度跨司法辖区评估(T3)。本文还提出Grounded Joint指标,该指标联合评估答案正确性与法律来源依据,并提供统一评估以简化基准测试。对代表性大语言模型的大量实验表明,尽管当前模型常能正确回答法律问题,但在提供准确的跨司法辖区法律推理方面存在困难。本文希望CrossLex能推动未来基于来源的跨司法辖区法律推理研究。

英文摘要

Legal reasoning is inherently jurisdiction-dependent: the same facts can call for different legal rules and yield different conclusions across legal systems. Yet existing benchmarks rarely evaluate whether large language models (LLMs) can recognize such jurisdiction-specific variation, especially when identical fact patterns lead to divergent legal outcomes.We introduce CrossLex, a same-fact, legal-source-grounded benchmark for evaluating cross-jurisdictional legal reasoning in LLMs across three jurisdictions: China, California, and Germany. Built from authoritative legal sources, CrossLex aligns 55 legal issues spanning contract, consumer, criminal, family, and labor law, and constructs jurisdiction-aligned questions paired with answers and supporting citations. In total, CrossLex contains 6,149 instances organized into 385 fact groups, with all legal issues, answers, and cited authorities reviewed by legal professionals.To disentangle basic legal knowledge from cross-jurisdictional reasoning, CrossLex defines three complementary tasks: single-jurisdiction reasoning (T1), joint cross-jurisdictional comparison (T2), and fine-grained cross-jurisdictional evaluation (T3). We further propose Grounded Joint, a metric that jointly assesses answer correctness and legal-source grounding, and provide a unified evaluation for streamlined benchmarking. Extensive experiments on representative LLMs show that, although current models can often answer legal questions correctly, they struggle to provide accurate cross-jurisdictional legal citations.We hope that CrossLex will facilitate future research on source-grounded cross-jurisdictional legal reasoning.

↑