arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

构建用于锚定与弃权(不执行)的法律奖励模型

Building Legal Reward Models for Grounding and Abstention

Rilton Franzone, Valentin Noël, Puyu Wang, Philip Torr, Fabio J. Fehr

arXiv 2609.14739首次发表:更新:

发表机构

University of Oxford; Devoteam(牛津大学; Devoteam)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出将法律问答数据转化为上下文偏好数据,构建LegalRewardBench基准,发现长度平衡增强和跨辖区迁移能显著提升法律RAG中奖励模型的锚定评估性能。

AI 中文摘要

大型语言模型越来越多地被用于法律等高风险领域,在这些领域中,系统必须将其推理锚定在检索到的证据上,并在证据不足时弃权(不执行)。然而,现有的奖励模型主要针对一般偏好而非上下文锚定进行优化,这限制了它们在检索增强生成(RAG)设置中评估这些行为的能力。我们引入了一个框架,将现有的法律问答数据集转换为上下文偏好数据,并利用该框架构建了LegalRewardBench(LRB),这是一个在嘈杂和不足的检索条件下评估锚定法律生成的基准。在一般和法律上下文评估中,我们发现上下文DPO改善了锚定评估,但性能对偏好数据的构建敏感。长度平衡增强显著改善了锚定法律评估,最强的配置结合了长度平衡的法律和一般上下文偏好数据,与基线相比性能提升了高达+25.6个百分点。我们进一步发现了跨司法管辖区迁移的证据:主要在维多利亚州刑法数据上进行上下文优化的模型,在外部美国法律基准上改善了锚定评估,包括在\ extsc{Housing Statute QA}上提升了+16.2个百分点。总之,这些结果为在检索增强设置中构建和评估锚定法律奖励模型提供了可复现的基础。

英文摘要

Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is insufficient. However, existing reward models are largely optimised for general preferences rather than contextual grounding, limiting their ability to evaluate these behaviours in retrieval-augmented generation (RAG) settings. We introduce a framework for transforming existing legal QA datasets into contextual preference data and use it to construct LegalRewardBench (LRB), a benchmark for evaluating grounded legal generation under noisy and insufficient retrieval conditions. Across general and legal contextual evaluation, we find that contextual DPO improves grounded evaluation, but performance is sensitive to preference-data construction. Length-balanced augmentation substantially improves grounded legal evaluation, with the strongest configuration combining length-balanced legal and general contextual preference data and improving performance by up to $\mathbf{+25.6}$pp over baseline. We further find evidence of cross-jurisdiction transfer: models contextually refined primarily on Victorian criminal-law data improve grounded evaluation on external US legal benchmarks, including a $\mathbf{+16.2}$pp improvement on \textsc{Housing Statute QA}. Together, these results provide a reproducible foundation for constructing and evaluating grounded legal reward models in retrieval-augmented settings.

CommentsPublished at ICML 2026 AI4Law Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑