arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IntLawNER:国际法中的命名实体识别数据集与基准

IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law

Genis Skura, Roland Bouffanais, Didier Wernli

arXiv 2609.22529首次发表:更新:

发表机构

University of Geneva(日内瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对国际法文本缺乏NER资源的问题,本文提出IntLawNER数据集与基准,通过混合算法-智能体流水线构建,并验证少样本提示可显著提升LLM在特定实体类型上的识别性能。

AI 中文摘要

国际法提供了国家协调行动、规范武装冲突和保护人权的规范框架,然而其文本至今缺乏词元级别的命名实体识别(NER)资源。我们推出了IntLawNER,一个针对国际法编纂来源的NER数据集与基准,涵盖来自国际法院(ICJ)判决、联合国安理会决议和欧洲人权法院(ECtHR)判决的2,987条黄金标注句子和8,094个实体跨度,并标注了七种特定于机构的实体类型。我们通过一种成本高效的混合算法-智能体流水线构建了IntLawNER,该流水线通过候选检索、基于LLM的审查和人工审核,将468k条源句子缩减为紧凑的标注集,其中89.6%的黄金跨度未经修改地从银层接受。然而,银层到黄金层的分析揭示,在特定领域的NER中,人机聚合一致性指标可能具有误导性:在边界匹配跨度上Cohen's kappa=0.964,但当包含缺失实体、边界错误和标签修正时,宏F1仅为0.753。基准测试表明,零样本基于跨度的GLiNER在依赖机构功能而非表面形式的实体类型上表现崩溃(0.243微F1),而微调后的变换器在稀有标签上表现挣扎。精心挑选的展示标签对比的少样本示例提升了所有LLM相对于零样本提示的性能,其中Claude Opus 4.6达到了最佳得分0.873微F1。我们发布IntLawNER作为提取国际法律文本中引用的基准和可复用资源。

英文摘要

International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its texts remain without token-level named entity recognition (NER) resources. We introduce IntLawNER, a NER dataset and benchmark for codified sources of international law, covering 2,987 gold-annotated sentences and 8,094 entity spans from International Court of Justice (ICJ) decisions, UN Security Council resolutions, and European Court of Human Rights (ECtHR) judgments, annotated with seven institution-specific entity types. We construct IntLawNER with a cost-effective hybrid algorithmic-agentic pipeline that reduces 468k source sentences to a compact annotation set through candidate retrieval, LLM-based vetting, and human review, with 89.6% of gold spans accepted unchanged from the silver layer. However, the silver-to-gold analysis reveals that human-machine aggregate agreement metrics can be misleading in domain-specific NER: Cohen's kappa=0.964 on boundary-matched spans masks a macro-F1 of 0.753 when missing entities, boundary errors, and label corrections are included. The benchmark shows that zero-shot span-based GLiNER collapses on entity types dependent on institutional function rather than surface form (0.243 micro-F1), while fine-tuned transformers struggle on rare labels. Carefully selected few-shot examples that demonstrate label contrasts improve every LLM over zero-shot prompting, with Claude Opus 4.6 reaching the best score of 0.873 micro-F1. We release IntLawNER as a benchmark and reusable resource for extracting references in international legal texts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑