arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从流程到证据:计算如何为对法律人工智能的适当依赖提供依据

From Process to Evidence: How Computing Can Ground Appropriate Reliance on Legal AI

James Bryan Williams

arXiv 2607.28869首次发表:更新:

AI 中文总结

本文针对法律AI应用中依赖校准缺乏证据、以流程替代证据的问题,将法律职责映射到HCI概念,提出相关要求与研究议程,助力缩小司法差距。

AI 中文摘要

律师和自行出庭的诉讼当事人已在使用人工智能(AI)起草法律文件,法院也出台了相关规则。在涉及AI幻觉的1500多起案件后,律师被要求对AI辅助提交的文件进行仔细的独立审查。履行这些职责需要人机交互(HCI)文献中所说的“适当依赖”,而要校准这种依赖,必须掌握这些工具在法律工作中出错的频率、严重程度以及可检测性的证据,现有研究几乎未描述这三个方面。我们分析了纽约法院系统的官方记录,这些记录反复要求提供不存在的证据(如错误率、禁用清单),却转而援引流程,包括培训要求、核对清单和未校准的人工审查,负担最落在最无力承担的群体身上:法律援助机构被告知要跟踪自身错误率,法官则只能自行制定测试。本文作出四项贡献:(1)将法律职责映射到HCI概念;(2)从官方记录中提炼出的一系列要求;(3)对法律系统以流程替代证据的分析;(4)为计算领域制定的研究议程,包括任务分类法、共享错误指标、维护的基准以及针对私有数据评估的测试工具。计算界必须提供司法系统所缺失的内容,此举有助于缩小而非扩大司法差距。

英文摘要

Lawyers and self-represented litigants are already using artificial intelligence (AI) to draft legal documents, and courts are responding with rules. After more than 1,500 cases involving AI hallucinations, lawyers have been instructed to perform careful, independent review of AI-assisted filings. Discharging these duties requires what the human-computer interaction (HCI) literature calls ``appropriate reliance,'' which cannot be calibrated without evidence on how often, how badly, and how detectably these tools fail at legal work. Existing research barely describes any of the three. We analyze the official record of the New York court system. The documents repeatedly call for evidence that does not exist (e.g., error rates, do-not-use lists). In its place they invoke procedure, including training mandates, checklists, and uncalibrated human review. The burden falls hardest on those least equipped to bear it: legal aid programs are told to track their own error rates, and judges are left to improvise their own tests. The paper makes four contributions: (1) a mapping from the legal duties to concepts in HCI; (2) a set of requirements elicited from the official record; (3) an analysis of how the legal system substitutes process for evidence; and (4) a research agenda for computing, including task taxonomies, shared error metrics, maintained benchmarks, and test harnesses for evaluations on private data. The computing community must supply what the justice system lacks; in doing so, it can help close, rather than widen, the justice gap.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑