arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

法律研究基准:衡量长时程法律研究智能体的端到端可靠性

Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents

Katrina Drozdov, Oliver Chen, Langston Nashold, Rayan Krishnan

arXiv 2610.00609首次发表:更新:

发表机构

Vals AI(Vals AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对法律研究智能体可靠性不足的问题,提出含413个专家问题的法律研究基准LRB,通过全通过评分与来源验证评估13个前沿模型,最强模型正确率仅42.9%。

AI 中文摘要

法律研究是一项核心且耗时的法律工作流程。律师必须识别具有约束力的权威依据,核实其仍然有效,协调法规与判例,并综合得出有根据的答案。语言模型智能体天然适合这种检索密集型工作流程,即使将其部分自动化也将具有重要价值。但这一价值取决于可靠性:一个缺失的权威依据、过时的引注或错误的法律结论,都可能使原本看似合理的答案变得不可用。我们提出了法律研究基准(Legal Research Bench, LRB),这是一个由专家撰写的413个开放式美国法律研究问题组成的基准,每个问题都配有标准答案、支持性权威依据和二元评分标准。我们在一个配备网络搜索、判例法搜索、页面解析和检索工具的实验框架中评估了十三个前沿模型。我们通过带来源验证的全通过评分来对智能体响应进行评分,即只有当每个必需标准都得到满足且其引用的权威依据得到验证时,响应才算正确。我们还针对专家律师验证了LLM评判器,确保基准分数与律师判断保持一致。智能体距离可靠仍有很大差距:在我们测试的模型中,最强的Claude Opus 4.8在42.9%的问题上完全正确。性能也因任务设置而有显著差异:全通过率在不同法律领域有所不同,并且在需要协调冲突权威依据的问题上更低。跨模型来看,更多的轮次、工具调用和推理成本并不能预测更高的准确性。

英文摘要

Legal research is a core and time-consuming legal workflow. Lawyers must identify controlling authority, verify that it remains valid, reconcile statutes and cases, and synthesize a grounded answer. Language model agents are a natural fit for this retrieval-intensive workflow, and automating even part of it would be valuable. But that value depends on reliability: a single missing authority, stale citation, or wrong legal conclusion can make an otherwise plausible answer unusable. We introduce \textbf{Legal Research Bench} (LRB), a benchmark of 413 open-ended U.S. legal research questions written by experts, each paired with a gold answer, supporting authorities, and a binary grading rubric. We evaluate thirteen frontier models in a harness with web search, case-law search, page parsing, and retrieval tools. We score agent responses through all-pass grading with source verification, where a response is correct only if every required criterion is satisfied and its cited authorities verify. We also validate the LLM judge against expert attorneys ensuring that benchmark scores track attorney judgment. Agents remain far from reliable: among the models we tested, the strongest, Claude Opus 4.8, is fully correct on 42.9\% of questions. Performance also varies substantially by task setting: all-pass rates differ across areas of law and are lower on questions requiring reconciliation of conflicting authorities. Across models, more turns, tool calls, and inference cost do not predict higher accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑