发表机构
Universidad de los Andes(安第斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对哥伦比亚法律体系构建专家验证基准,评估15个LLM,发现自由文本正确性低且与相关性分离,需专家监督,引用权威来源可提升可靠性。
AI 中文摘要
大型语言模型(LLM)越来越多地被用于支持法律实践、教育和研究,然而它们在美国以外的国家法律体系中的可靠性在很大程度上仍未被记录。我们引入了一个经过专家验证的基准,用于评估LLM在哥伦比亚法律体系上的可靠性。该基准包含1,042个项目,涵盖十个法律领域和三种问题格式(封闭式多项选择、半开放式和开放式IRAC),通过人工参与流程构建,并经过多阶段专家审查。我们使用与格式相匹配的指标评估了15个当代专有和开放权重模型。封闭式问题的准确率范围广泛,从0.905(Gemini 3.1 Pro)到0.577,但在自由文本法律答案上,任何模型的事实正确性均未超过0.45(在0-1量表上)。我们发现答案相关性与正确性之间存在分离(Spearman rho = -0.46):模型可靠地听起来有回应,却经常出错,这种模式对非专家用户尤其令人担忧。封闭式问题准确率与自由文本正确性之间存在强秩相关(rho = 0.94),因此廉价的多项选择筛选可以预测模型排名,但会高估绝对可靠性。一个独立的基于评分标准的LLM评判者和盲法人类专家评分均重现了自由文本排名(rho >= 0.88)。评判者进一步揭示,模型引用的规范中只有约一半是正确的;其余是错误的或不存在的。可靠性随法律领域系统性变化,并随问题复杂性呈倒U型分布。我们的结果表明,当前LLM在哥伦比亚法律任务中需要专家监督,而将答案基于权威来源是提高可靠性的有前景的途径。我们发布了基准构建流程,以支持可复现的评估。
英文摘要
Large language models (LLMs) are increasingly used to support legal practice, education, and research, yet their reliability in national legal systems outside the United States remains largely undocumented. We introduce an expert-validated benchmark for evaluating LLM reliability on the Colombian legal system. The benchmark comprises 1,042 items spanning ten areas of law and three question formats (closed multiple-choice, semi-open, and open-ended IRAC), built through a human-in-the-loop pipeline with multi-stage expert review. We evaluate 15 contemporary proprietary and open-weight models with format-appropriate metrics. Accuracy on closed questions ranges widely, from 0.905 (Gemini 3.1 Pro) to 0.577, but on free-text legal answers factual correctness never exceeds 0.45 (on a 0-1 scale) for any model. We find a dissociation between answer relevancy and correctness (Spearman rho = -0.46): models reliably sound responsive while frequently being wrong, a pattern of particular concern for non-expert users. Closed-question accuracy and free-text correctness are strongly rank-correlated (rho = 0.94), so cheap multiple-choice screening predicts model ranking but overstates absolute reliability. An independent rubric-based LLM judge and blind human expert scoring both reproduce the free-text ranking (rho >= 0.88). The judge further reveals that only about half of the norms models cite are correct; the rest are wrong or non-existent. Reliability varies systematically by legal area and follows an inverted-U across question complexity. Our results indicate that current LLMs require expert supervision for Colombian legal tasks, and that grounding answers in authoritative sources is a promising path to higher reliability. We release the benchmark construction pipeline to support reproducible evaluation.
Comments38 pages, 23 figures, 8 tables