arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10293cs.CLcs.AIcs.IR

GANDR:用于可验证法律答案生成的主张审计

GANDR: Claim Auditing for Verifiable Legal Answer Generation

Chen Qian, Yimeng Wang, Yu Chen, Lingfei Wu, Andreas Stathopoulos

首次发表
浏览论文内容

中文总结 AI 辅助

GANDR通过双智能体系统(起草器与评审器)对法律答案逐项主张进行审计,在185项基准上以70.8%严格准确率领先最强基线11.3个百分点,实现可验证的生成。

中文摘要 AI 辅助

在诸如法律实践等高风险领域,语言模型的答案只有在读者能够对照系统引用的来源验证每项主张时才有用。当前的基于检索的生成流程对答案整体评分,因此正确的结论可能建立在捏造或松散匹配的引用之上,却仍能获得高分。弥合这一差距既需要一个为逐项主张验证而构建的系统,也需要一种衡量该能力的评估方法。我们提出了GANDR(Grounded ANswer DRafter,基于检索的答案起草器),这是一个双智能体系统,其中起草器以结构化的法律推理格式撰写答案,而一个独立的评审器,与人工验证者具有相同的视角,对每项主张与其引用的来源进行审计,并在每轮输出逐项主张的审计轨迹。我们为其配备了一个严格正确性标准,要求每个引用都能解析到检索器返回的段落。在一个包含185个条目的法律基准测试中,所有六个系统共享同一个主干模型、同一个检索面以及同一条引用指令,GANDR在所有主要指标上均排名第一,达到了70.8%的严格准确率,并以11.3个百分点(p<0.01)的优势领先于最强基线。将协议锚定的提交规则恢复为默认设置会使严格准确率降低22.7个百分点,并且在另外三个主干模型上,严格准确率的领先优势保持为正,介于+3.2到+6.5个百分点之间。这一领先优势归因于起草器的配置和协议锚定的提交,而非改写。与两位受过法律训练的标注者相比,该审计作为二元检测器,在识别支持不足的主张时达到了F1分数0.84,而其四分类判定标签仅具有弱一致性,且仅供参考。代码可根据要求提供。

英文摘要

In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the source the system cites. Current grounded-generation pipelines score the answer as a whole, so a correct conclusion can rest on fabricated or loosely matched citations and still score well. Closing this gap requires both a system built for per-claim verification and an evaluation that measures it. We introduce GANDR (Grounded ANswer DRafter), a two-agent system in which a Drafter writes an answer in a structured legal-reasoning format and a separate Critic, with the same view as a human verifier, audits each claim against its cited source and emits a per-claim audit trace on every round. We pair it with a strict correctness criterion requiring every citation to resolve to a passage the retriever returned. On a 185-item legal benchmark where all six systems share one backbone, one retrieval surface, and one citation instruction, GANDR ranks first on every primary metric, reaching 70.8% strict accuracy and leading the strongest baseline by 11.3 points (p<0.01). Reverting the protocol-anchored commit rule lowers strict accuracy by 22.7 points, and the strict lead stays positive on three further backbones, at +3.2 to +6.5 points. This lead traces to the Drafter configuration and the protocol-anchored commit, not to rewriting. Against two law-trained annotators the audit flags under-supported claims at F1 0.84 as a binary detector, while its four-way verdict labels agree only weakly and are advisory. Code is available upon request.

发表机构

  • William & Mary(威廉与玛丽学院)
  • Anytime AI

机构由 AI 辅助整理,请以论文原文为准。

↑