发表机构
Lexsi Labs(雷克西实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过反事实审计发现,大型语言模型提及法律权威的表现与其判决随权威变化的一致性存在巨大差距,专用模型也无法缩小差距,且模型易受对抗性操纵,这对法律AI的应用具有重要意义。
AI 中文摘要
大型语言模型越来越多地通过提及判决背后的法规或先例来为法律决策提供理由,这被视为决策遵循该法规或先例的证据。我们对此进行了直接测试:在保持案件事实不变的情况下,将提及的法律权威替换为不相关的权威,并从模型的隐藏状态中解码其最终判决。在7个开放权重模型(8B至70B)和4个涵盖司法与合同推理的基准测试中,当被明确要求通过提及管辖权威来为判决提供理由时,模型在66.7%至100%的生成内容中提及了正确的权威,但当权威发生变化时,判决随之改变的一致性要低得多:CaseHOLD数据集上为0.0%至21.7%,ECHR和SCOTUS数据集上为30.0%至76.7%,ContractNLI数据集上为43.3%至50.0%。模型规模或专用法律推理模型(尽力的LoRA复现;第6节)均未缩小这一差距。对5个核心模型的红队评估发现,对隐藏在案件事实中的对抗性指令的遵从度(73.3%至96.4%)远高于判决切换敏感性,且在所有模型排名中均无例外。因此,提及法律权威是判决对其依赖程度的糟糕替代指标,同时相同判决仍易受对抗性操纵。两项发现均通过排除提示措辞噪声和混淆采样的检查得到复现,且直接关系到将生成的法律解释用作合规或审计人工制品的应用场景。
英文摘要
Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from its hidden states. Across seven open-weight models (8B-70B) and four benchmarks spanning judicial and contractual reasoning, when explicitly required to justify a verdict by naming the governing authority, models name the correct one in 66.7%-100% of generations, while the verdict changing when the authority changes is far less consistent: 0.0%-21.7% on CaseHOLD, 30.0%-76.7% on ECHR and SCOTUS, and 43.3%-50.0% on ContractNLI. Neither scale nor a purpose-built legal-reasoning model (a best-effort LoRA reproduction; Section 6) closes this gap. A red-teaming evaluation on five core models finds compliance with an adversarial instruction hidden in the case facts (73.3%-96.4%) exceeds verdict-swap sensitivity by a wide margin, holding without exception across model rankings. Naming a legal authority is thus a poor proxy for a verdict's dependence on it, while the same verdict remains separately vulnerable to adversarial manipulation. Both findings replicate across checks ruling out prompt-wording noise and confounded sampling, and bear directly on the use of generated legal explanations as compliance or audit artefacts.