大语言模型能否以具有法律意义的方式进行推理?针对欧洲人权法院案件的小规模研究
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases
浏览论文内容
中文总结 AI 辅助
本研究以欧洲人权法院案件为测试平台,评估OpenAI GPT 5.4的法律推理表现,发现其推理质量不足、LLM评估不可替代人工,且专家提示未提升预测准确率。
中文摘要 AI 辅助
推理已成为当代大语言模型(LLM)的标准技术与特征,但其在法律导向的高要求任务(如法律案件预测)中的应用及质量仍未得到充分探索。本研究以欧洲人权法院(ECtHR)的法律案件为测试平台,探究LLM在法律案件预测语境下的推理表现。我们通过探索不同提示策略(这些策略或多或少暗示了ECtHR判例中具有法律意义的推理的标准),评估了近期顶级LLM OpenAI GPT 5.4。我们呈现了通过人工与LLM评估模型响应所得的研究发现:所评估模型的法律推理得分远未达到理想水平,其生成的分析在结构上完整但内容上较浅;LLM作为评判者(LLM-as-a-Judge)的评估者内部一致,但与我们训练的标注者仅存在弱一致性,即LLM评估可靠但不能有效替代人工评估。总体而言,经专家精心设计的提示能带来更全面的推理,但与其他所评估的设置相比,并未产生更准确的预测。基于这些发现,我们呼吁学界不要仅依赖基于LLM的自动化评估,且避免将任务准确率作为推理质量的合适替代指标。
英文摘要
Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain under explored. We investigate how LLMs reason in the context of legal case forecasting, using legal cases from the European Court of Human Rights (ECtHR) as a testbed. We evaluate OpenAI GPT 5.4, a recent top-tier LLM, by exploring alternative prompting strategies that are more or less suggestive of what counts as legally meaningful reasoning in the context of ECtHR jurisprudence. We present our findings derived from assessing the model's responses with both human and LLM evaluation. We find that the examined model scores far from ideal in legal reasoning, the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet align only weakly with our trained annotators, i.e., reliable but not a valid substitute for human evaluation. Overall, the expert-curated prompt leads to more comprehensive reasoning, which does not result in more accurate predictions compared to the other examined settings. Based on our findings, we urge the community not to rely solely on automated LLM-based evaluation and to avoid using task accuracy as an appropriate proxy for reasoning quality.
发表机构
- University of Copenhagen(哥本哈根大学)
- Faculty of Law, University of Copenhagen(哥本哈根大学法学院)
- Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。