arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39071cs.CL

LexReward:面向法律语言模型的分类学驱动奖励框架

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models

  • Peking University(北京大学)
  • Northeastern University(东北大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Yida Cai, Xin Dai, Bingxiang He, Huiyuan Xie, Yuxiao Ye, Zhenghao Liu, Yang Bai, Zhiyuan Liu

中文总结 AI 辅助

LexReward提出分类学驱动的奖励框架,从风格、要素、链条三维度评估法律回答质量,构建偏好数据训练DPO和奖励模型,实验证明其有效提升各维度性能。

中文摘要 AI 辅助

法律语言模型需要能够捕捉回答正确性以及法律回答多维质量的奖励信号。然而,现有的奖励方法往往依赖粗粒度的整体判断,提供的领域特异性和可解释性有限。我们提出LexReward,一个用于法律奖励建模的分类学驱动框架。LexReward从三个互补维度刻画法律回答质量:风格(Style),涵盖词汇和句法质量;要素(Element),评估法律主体、事实、法条和判决;链条(Chain),评估法律推理的顺序、完整性、正确性和非冗余性。针对每个维度,我们制定了明确评价标准和质量等级的评分细则。由此产生的奖励被用于构建直接偏好优化(DPO)和奖励模型训练的成对偏好数据。实验表明,基于评分细则的奖励能够可靠地区分不同质量的法律回答,且基于偏好数据的DPO训练在三个维度上均提升了性能。学习得到的奖励模型LexRM也支持有效的下游优化:每个维度的奖励模型通过强化学习在对应维度上提升策略性能,且奖励时无需参考答案。逐维度分析进一步支持了所提分类学和奖励构建的有效性。

英文摘要

Legal language models require reward signals that capture not only answer correctness but also the multidimensional quality of legal responses. Existing reward methods, however, often rely on coarse-grained holistic judgments, providing limited domain specificity and interpretability. We introduce LexReward, a taxonomy-driven framework for legal reward modeling. LexReward characterizes legal response quality along three complementary dimensions: Style, covering lexical and syntactic quality; Element, assessing legal subjects, facts, statutes, and decisions; and Chain, evaluating the order, completeness, correctness, and non-redundancy of legal reasoning. For each dimension, we develop rubrics that specify evaluation criteria and quality levels. The resulting rewards are used to construct pairwise preference data for Direct Preference Optimization (DPO) and reward-model training. Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions. The learned reward models, LexRM, also support effective downstream optimization: each dimension-specific reward model improves policy performance in its corresponding dimension through reinforcement learning, without requiring reference answers at reward time. Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.

↑