arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27678cs.CL

相同分数,不同决策:评估 JEV 与语言模型在法律文档理解中的表现

Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding

Fan Zhang, Yankai Chen, Zhuohan Xie, Yixi Zhou, Sijia Peng, Lei Fan, Xinhua Ji, Cunyuan Zheng, Huangyong Shan, Philip S. Yu, Xue Liu, Yu Chen, Preslav Nakov, Songwei He

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在合同推断任务上对比Jev与九个语言模型,发现低成本与快响应伴随较低基线准确率,且准确率排名与各条件下正确性排名不一致,强调需综合评估成本、速度及个体判断稳定性。

中文摘要 AI 辅助

合同推断需要对同一文档做出多项判断,但总体准确率可能掩盖个体决策的变化。重复的一致性也不足以说明问题:模型可能始终返回错误答案。本文在 ContractNLI 上将 Jev 与九个语言模型进行比较,评估推理成本、响应时间、平均正确率以及重复请求条件下的正确性。受控比较在保持合同与目标判断不变的前提下,改变假设可见性、请求输出和输出顺序。在所评估的配置中,Jev 具有最低成本和最短中位响应时间,而托管语言模型则达到更高的基线准确率。按基线准确率的排名与按每个条件和重复下的正确性的排名不同,尽管后者中的微小差异并不能确立普遍的稳定性优势。开发诊断进一步揭示了补偿性修正与回退,以及持续存在的错误。这些发现促使在评估成本与响应时间的同时,关注个体判断在请求配置变化时是否保持正确。代码:https://github.com/ZF-Utokyo/Jev-Benchmark

英文摘要

Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating inference cost, response time, average correctness, and correctness across repeated request conditions. Controlled comparisons vary hypothesis visibility, requested outputs, and output order while keeping the contract and target judgment fixed. Jev has the lowest cost and median response time among the evaluated configurations, while hosted language models achieve higher baseline accuracy. Rankings by baseline accuracy differ from rankings by correctness across every condition and repeat, although small differences in the latter do not establish a general stability advantage. Development diagnostics further reveal compensating corrections and regressions, as well as persistent errors. These findings motivate evaluating cost and response time alongside whether individual judgments remain correct as the request configuration changes. Code: https://github.com/ZF-Utokyo/Jev-Benchmark

发表机构

  • The University of Tokyo(东京大学)
  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)
  • McGill University(麦吉尔大学)
  • Hong Kong Baptist University(香港浸会大学)
  • Fudan University(复旦大学)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • UCloud(优刻得)
  • Columbia University(哥伦比亚大学)
  • The University of Hong Kong(香港大学)
  • University of Illinois Chicago(伊利诺伊大学芝加哥分校)
  • Quantell Capital

机构由 AI 辅助整理,请以论文原文为准。

↑