arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20391cs.CL

ImmigrationReason:面向法律推理研究的美国移民上诉结构化数据集

ImmigrationReason: A Structured Dataset of U.S. Immigration Appeals for Legal Reasoning Research

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Amirhossein Afsharrad, Seyed Shahabeddin Mousavi

AI总结:

本研究推出ImmigrationReason大规模结构化数据集,涵盖美国移民上诉裁决相关数据,可支撑法律推理领域的结果预测、错误分析及高风险监管智能体设计等研究方向。

AI中文摘要:

大多数法律自然语言处理(NLP)资源均取自联邦判例法,且侧重粗分类,而占政府决策绝大多数的行政裁决领域基本未被涉及。我们推出ImmigrationReason,这是一个大规模结构化数据集,源自美国公民及移民服务局(USCIS)行政上诉办公室(AAO)2005年至2026年间的12375项非先例裁决。每条记录涵盖适用法律框架、基于五类标签的各标准证据充分性判定结果、裁决者批评的逐字引述、所有引用文献及最终处置结果,以及由Claude转录的高质量源文本。提取质量通过三轮流程验证,该流程结合两种独立模态并使用Opus 4.7的比较提示进行判定,且由领域专家对500条记录样本进行了核查。该数据集记录了近9000个AAO认定的法律错误逐字实例,涵盖自然法律制度转变(2016年Dhanasar规则变更),覆盖21年的裁决数据。我们对该数据集进行了详细分析,并概述了它所能支撑的研究方向,涵盖结果预测、裁决者错误分析以及高风险监管领域的智能体设计等。

英文摘要:

Most legal NLP resources draw from federal case law and focus on coarse classification, leaving administrative adjudication, where the vast majority of government decisions occur, essentially unaddressed. We introduce ImmigrationReason, a large-scale structured dataset derived from 12,375 non-precedent decisions of the U.S. Citizenship and Immigration Services (USCIS) Administrative Appeals Office (AAO) spanning 2005 to 2026. Each record captures the applicable legal framework, per-criterion evidence-sufficiency findings under a five-category label, verbatim adjudicator-criticism quotes, all citations, and final dispositions, alongside high-quality Claude-transcribed source text. Extraction quality is validated through a three-pass pipeline combining two independent modalities with comparison-prompt adjudication by Opus 4.7, and verified by domain experts on a 500-record sample. The dataset documents nearly 9,000 verbatim instances of AAO-identified legal errors, spans a natural legal-regime transition (the 2016 Dhanasar rule change), and covers 21 years of adjudication. We analyze the dataset in detail and outline research directions it enables, from outcome prediction and adjudicator-error analysis to agent design for high-stakes regulatory domains.

↑