发表机构
University of Copenhagen(哥本哈根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有法律IR数据集缺乏段落级引用标注、存在数据泄露等局限,构建源自CJEU判决的LegalPincite数据集,支持多级别法律IR方法的开发与严格评估。
AI 中文摘要
法律信息检索(IR)的一项常见任务是从判例法集合中查找相关法律来源。法律实践通常需要精准引用(pincite)到具体判例段落,但大多数现有公开法律IR数据集缺乏段落级引用标注。不过,包含此类信息的公开数据集存在查询文本数据泄露问题,且未纳入既非引用也非被引用的段落,形成不现实、过于简化的检索设置,可能导致性能评估虚高。为解决这些局限,我们构建了一个源自欧洲联盟法院(CJEU)判决的大规模法律IR数据集。该数据集包含:(i)移除了引用信息的掩码判例/段落查询;(ii)涵盖所有段落的语料库;(iii)判例级和段落级的真实引用,部分经人类专家验证。本数据集支持在多个查询-文档级别(判例到判例、段落到判例、段落到段落检索)开发和严格评估法律IR方法。数据集链接:this https URL
英文摘要
A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While legal practice often requires pinpoint citations (pincites) to specific case paragraphs, most existing public legal IR datasets lack paragraph-level citation annotations. Yet, publicly available datasets with such information contain data leakage in the query text and exclude paragraphs that are neither citing nor cited from the corpora, creating an unrealistic and oversimplified retrieval setting, potentially leading to inflated performance. To address these limitations, we contribute a large-scale legal IR dataset constructed from Court of Justice of the European Union (CJEU) judgments. The dataset contains: (i) masked case/paragraph queries, with removed citation information; (ii) a corpus that includes all paragraphs; and (iii) case- and paragraph-level ground truth citations, with partial human expert validation. Our dataset supports both the development and rigorous evaluation of legal IR methods, at multiple query-document levels (case-to-case, paragraph-to-case, and paragraph-to-paragraph retrieval). Link to dataset and code: https://huggingface.co/datasets/theresiavr/legalpincite
CommentsAccepted for publication at the 8th Natural Legal Language Processing Workshop (NLLP 2026), co-located with EMNLP 2026