PROSLEX:印度司法领域专家标注的法律条文预测新型数据集
PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary
浏览论文内容
中文总结 AI 辅助
本文针对法律条文预测研究中缺乏可解释性的问题,构建印度司法领域专家标注的PROSLEX数据集,评估多种提示策略,为可解释法律AI提供基准。
中文摘要 AI 辅助
法律条文预测(Legal Statute Prediction, LSP)是指根据法律文件中的事实描述自动识别相关法律条文,通常被视为自然语言处理与信息检索研究中的多标签分类任务。尽管近期研究已开始将大语言模型(Large Language Models, LLMs)应用于条文预测,但现有方法主要关注准确率指标,未解决法律推理这一关键需求——司法场景中决策必须具备可解释性与正当性。为填补该研究空白,本文提出PROSLEX(PRediction Of Statutes and LEgal eXplanation,即条文预测与法律解释),这是一个包含1623份印度语境下专家标注法律文件的综合数据集,每份文件均配有条文预测结果与详细解释,总计7450条解释,涵盖了潜在的法律推理过程。利用该数据集,本文系统评估了各类提示策略,包括零样本、少样本、思维链及树状思维方法,以生成条文预测结果及其对应的法律理由。本文的评估框架不仅衡量预测性能,还评估生成解释的连贯性与法律有效性,使PROSLEX成为开发可解释AI系统的基准,该系统可支持法律从业者并推动可解释法律自然语言处理研究。为确保可复现性,本文已将PROSLEX数据集与模型代码发布在GitHub上:this https URL。
英文摘要
Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research. While recent advances have begun incorporating Large Language Models (LLMs) for statute prediction, current approaches primarily focus on accuracy metrics without addressing the critical need for legal reasoning, a fundamental requirement in judicial contexts where decisions must be explainable and justifiable. To address this research gap, we present PROSLEX (PRediction Of Statutes and LEgal eXplanation), a comprehensive dataset comprising 1,623 expert-annotated legal documents from the Indian context. Each document is paired with statute predictions and detailed explanations, totaling 7,450 explanations, capturing the underlying legal reasoning. Using this dataset, we systematically evaluate various prompting strategies, including zero-shot, few-shot, chain-of-thought, and tree-of-thoughts approaches, to generate both statute predictions and their corresponding legal rationales. Our evaluation framework measures not only predictive performance but also the coherence and legal validity of generated explanations, positioning PROSLEX as a benchmark for developing explainable AI systems that can support legal practitioners while advancing research in interpretable legal NLP. To ensure reproducibility, we have made our PROSLEX dataset and model code available on GitHub: https://github.com/subinay494/Legal_Statute_Prediction_Explanation.
发表机构
- IISER Kolkata(印度科学教育与研究学院加尔各答分校)
- Utrecht University(乌得勒支大学)
- IIT Kharagpur(印度理工学院克勒格布尔分校)
- University of Birmingham Dubai(伯明翰大学迪拜分校)
- WBNUJS(西孟加拉邦国家法律大学)
机构由 AI 辅助整理,请以论文原文为准。