发表机构
Epiq AI Labs; Cornell University(Epiq AI 实验室; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出法律调查智能体环境InvestigationWorlds,基于真实美国联邦案件构建多解读语料库,评估发现智能体难以区分法院采纳假设与替代假设。
AI 中文摘要
我们引入了InvestigationWorlds,一个用于法律调查的智能体环境。我们基于美国民事诉讼中一个未被充分利用的文书——即决判决动议(summary judgment motion)进行构建。该动议依赖于由真实证据展品组成的记录,并产生一个法院采纳的假设,该假设在决定动议时被视为事实依据。每个环境都基于从法院电子记录公共访问系统(PACER)检索到的真实美国联邦法院案件构建,并通过一个经律师验证的生成流水线进行增强,该流水线围绕原始记录合成带有角色标签的文档。由此产生的语料库允许多种连贯的事实解读,其中只有一种与法院采纳的假设相符。在100个案件上的评估中,我们发现智能体尽管检索到了相关证据,却常常固守错误的假设,难以将法院采纳的假设与替代假设区分开来。
英文摘要
We introduce InvestigationWorlds, an agentic environment for legal investigation. We build on an underused artifact of U.S. civil litigation: the summary judgment motion. This motion relies upon a record composed of real evidence exhibits, and results in a court-adopted hypothesis that is treated as ground truth for the purposes of deciding the motion. Each environment is built from a real U.S. Federal Court case retrieved from Public Access to Court Electronic Records (PACER) and augmented by an attorney-validated generation pipeline that synthesizes role-tagged documents around the original record. The resulting corpus admits multiple coherent factual readings, only one of which matches the court-adopted hypothesis. Evaluating on 100 cases, we find agents often commit to incorrect hypotheses despite retrieving relevant evidence, struggling to distinguish the court-adopted hypothesis from alternative hypotheses.
CommentsAccepted to NeurIPS 2026 Evaluations & Datasets