arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知晓形式,而非功能:自动审计法律基准中答案与权威的脱钩问题

Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks

Hsien-Jyh Liao

arXiv 2608.02621首次发表:更新:

发表机构

Ministry of Justice, Taiwan(台湾法务部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对法律基准中答案与权威脱钩问题,联合审计四款LLMs在台湾律师考试题上的表现,发现答案与引用行为可独立变化,提出答案与权威联合评估的方法。

AI 中文摘要

法律基准通常会对模型给出的最终答案进行评分,即便模型同时也会引用法律权威依据。本研究测试了答案正确性能否作为权威依据的替代指标。在未要求引用法规的普通推理提示下,四款大语言模型(LLMs)在238道台湾律师考试题目中自发产生了权威标记。由于每道题目都有经核实的管辖条款,我们自动联合审计了答案正确性与权威依据情况,发现两个维度会向两个方向脱钩:在刑法领域,24.0%至42.4%的有效回答答案正确但遗漏了权威依据,15.2%至21.7%的回答答案错误但引用了权威依据。单独的法规检索探测及宽松的引用弃权(不执行)干预进一步表明,答案与引用行为在输出层面可独立变化。由于这种不匹配并非由对抗性或不一致提示引发,仅基于答案的评分会将自然出现的遗漏权威依据的情况视为基准的完全成功。鉴于法规权威在结构上可提取且可外部验证,该失败可自动测量。初步的中国民法扩展研究也观察到未被要求引用时的权威标记,推动开展全面的跨司法辖区联合审计。因此,本文提出针对基于法规的法律基准,采用答案与权威联合评估的方法。

英文摘要

Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy for authority grounding. Under ordinary reasoning prompts that did not request statutory citations, four LLMs spontaneously produced authority markers across 238 Taiwan bar-examination items. Because each item has a verified governing provision, we automatically audit answer correctness and authority grounding jointly. The two dimensions dissociate in both directions. In criminal law, 24.0--42.4\% of valid responses were answer-correct but missed the gold authority, while 15.2--21.7\% were answer-incorrect but cited it. A separate statutory-retrieval probe and a permissive citation-abstention intervention further show that answer and citation behavior can move separately at the output level. Because this mismatch arises without adversarial or inconsistency-inducing prompting, answer-only scoring treats naturally occurring gold-authority misses as complete benchmark successes. Because statutory authority is structurally extractable and externally verifiable, the failure can be measured automatically. A preliminary PRC civil-law extension also observes citation-unrequested authority marking, motivating a full cross-jurisdictional joint audit. We therefore propose joint answer--authority evaluation for statute-grounded legal benchmarks.

Comments9 pages, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑