AI 中文总结
研究长法律文本检索难题,提出提取法律要点并嵌入索引进行检索的方法,在四个基准上评估,结果显示该方法对部分法学数据集有显著提升,对其他场景效果有别,证明法律要点在特定法学搜索中有用。
AI 中文摘要
法律文献检索具有挑战性,因为法院判决文本长且内容多样,相关法律论点可能只占一小部分。本文探讨从源文档中提取的简短且自包含的法律要点能否改进巴西法律文献的密集检索。我们提出了一个流程,先从文档中提取要点,用嵌入索引,检索要点级证据,再将检索到的要点聚合回文档级排名。在四个葡萄牙法律检索基准上评估该方法,报告了NDCG@10、MAP@10和MRR@10。要点检索显著改善了两个法学数据集,但在NormasTCU和BR - TaxQA上不如全文检索。结果表明法律要点对法学搜索有用,尤其是查询以法律论点形式提出时,但在其他法律检索场景中效果可能不同。
英文摘要
Legal retrieval over jurisprudential collections is challenging because court decisions are long, heterogeneous documents whose relevant legal thesis may occupy only a small portion of the text. This paper asks whether legal nuggets, defined as short and self-contained legal theses extracted from source documents, can improve dense retrieval over Brazilian legal collections. We propose a pipeline that extracts nuggets from each document, indexes them with embeddings, retrieves nugget-level evidence, and aggregates the retrieved nuggets back to document-level rankings. We evaluate this approach on four Portuguese legal retrieval benchmarks from the JUA ecosystem, reporting NDCG@10, MAP@10, and MRR@10. Nugget retrieval substantially improves the two jurisprudential datasets: on JUA-Juris, NDCG@10 increases from 0.10265 to 0.20461, and on JurisTCU from 0.20898 to 0.32696. However, it underperforms full-document retrieval on NormasTCU and BR-TaxQA, and an embedding-model ablation shows that strong domain-adapted retrievers can remain better in the full-document setting. The results demonstrate that legal nuggets can be useful for jurisprudence search, especially when queries are formulated as legal theses, but they may not transfer equally well to other legal retrieval scenarios.