arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14610cs.AI

大型语言模型何时适用错误法律?诊断大型语言模型在时间法律推理中的失败

When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning

发表机构东京科学大学 · 大阪大学 · 日本信息学研究所语言、语言学与媒体中心
另 2 家 · 查看机构详情
  • Institute of Science Tokyo(东京科学大学)
  • Osaka University(大阪大学)
  • NII LLMC(日本信息学研究所语言、语言学与媒体中心)
  • Nara Institute of Science and Technology(奈良科学技术研究所)
  • Center of Juris-Informatics, ROIS-DS(日本学术研究推进机构跨学科科学研究中心法学信息学中心)

机构由 AI 辅助整理,请以论文原文为准。

Yiqian Huang, Shuyuan Zheng, Qianying Liu, Shaowen Peng, Yuntao Kong, Kotaro Funakoshi, Chuan Xiao, Manabu Okumura, Yang Cao

首次发表
浏览论文内容

中文总结 AI 辅助

本文构建基准评估LLMs的时间适用法律判定能力,发现LLMs存在适用最新法律的偏差,该偏差源于强化学习塑造的显式推理导致推理路径多样性减少,通用推理能力更强的模型在时间法律推理中表现更差,为相关优化提供指导。

中文摘要 AI 辅助

法律判决预测(LJP)等法律推理任务需要识别管辖案件的时间上正确的法律版本——我们将这一能力称为时间适用法律判定。然而,大型语言模型(LLMs)能否可靠执行该任务仍未被探索。在本文中,我们构建了一个基准来评估LLMs在时间适用法律判定方面的表现,并系统研究它们在时间法律推理中失败的原因。我们的实验揭示了四个关键发现:第一,LLMs表现出强烈的适用最新颁布法律的偏差,无论与法律相关的事实发生在何时;第二,这种偏差并非源于无法理解法律具有时间范围,也非源于缺乏对历史法规的知识;第三,我们提供了行为证据,表明强化学习塑造的显式推理可能是关键机制:在提升通用推理能力的同时,它减少了推理路径的多样性,导致模型收敛于适用现行法律;第四,这产生了一种反常的负相关关系:通用推理能力更强的模型在时间法律推理上表现更差。我们的发现为未来提升LLMs在基于时间的法律推理中的表现提供了具体指导。

英文摘要

Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination. However, whether large language models (LLMs) can reliably perform this task remains unexplored. In this paper, we construct a benchmark to evaluate LLMs on temporal applicable-law determination, and systematically investigate why they fail at temporal legal reasoning. Our experiments reveal four key findings. First, LLMs exhibit a strong bias toward applying the most recently enacted law, regardless of when the legally relevant facts occurred. Second, this bias does not stem from an inability to understand that laws have temporal scope, nor from a lack of knowledge about historical statutes. Third, we provide behavioral evidence that reinforcement-learning-shaped explicit reasoning may be a key mechanism: while improving general reasoning ability, it reduces the diversity of reasoning paths, causing models to converge on applying the current law. Fourth, this produces a counterintuitive inverse relationship: models with stronger general reasoning ability tend to perform worse on temporal legal reasoning. Our findings offer concrete guidance for future work on improving LLM performance in temporally grounded legal reasoning.

↑