发表机构
National Institute of Technology, Tiruchirappalli; National Institute of Karnataka, Surathkal(国立技术学院特鲁奇拉帕利分校; 卡纳塔克国立学院苏拉萨尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对印度法律先例检索,提出结合修辞角色、查询分割与冻结验证融合的四阶段流水线,显著优于单一角色和全文检索,并揭示了LLM重排序器的评估偏差。
AI 中文摘要
在印度法律判决中检索先例案件时,问题在于区分相关的法律事实与单纯的词汇相似性,因为与查询判决共享法规的先例案件不一定相关。本文描述了对IL PCR(印度法律先例检索)任务三种连续设计的检索系统的实证评估。我们证明,修辞角色(事实、判决理由、先例、论点、法规)的组合优于单独使用各角色或进行全文检索,法规相似性不具有区分性,并且使用LLM生成的训练数据的法律蕴含重排序器在官方评估中的表现远差于其内部验证分数。两个排名器内部MRR高达0.97,但在官方评估中表现不佳(MRR低至0.14)。因此,我们提出了一种新设计,采用查询不相交分割和冻结验证融合。结果是四阶段流水线,Micro F1为0.2549,MRR为0.6081,nDCG@10为0.4187。
英文摘要
For retrieving prior cases in Indian legal judgments, the problem involves distinguishing relevant legal facts from mere lexical similarities because a prior case that shares a statute with the query judgment is not necessarily relevant. In this paper, we describe an empirical evaluation of three consecutive designs of retrieval systems for the IL PCR(Indian Legal Prior Case Retrieval) task. We demonstrate that the combination of the rhetorical roles (Fact, Ratio Of The Decision, Precedent, Argument, Statute) is better than either using individual roles or performing full text retrieval, that statute similarity is not discriminative, and that a legal entailment reranker with training data produced by an LLM is much worse in terms of official evaluation than its internal validation score. Two rankers, with excellent internal MRR up to 0.97, performed poorly in terms of official evaluation (MRR as low as 0.14). Thus, we propose a new design with query disjoint splitting and frozen validation fusion. The result is a four stage pipeline with Micro F1 0.2549, MRR 0.6081, and nDCG@10 0.4187.