arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GreekBarRetrieval:希腊成文法检索基准

GreekBarRetrieval: A Benchmark for Greek Statutory Retrieval

Ernest Beta, Odysseas S. Chlapanis, Dimitrios Galanis, Ion Androutsopoulos

arXiv 2608.18752首次发表:更新:

发表机构

Archimedes, Athena Research Center; Institute for Language and Speech Processing, Athena Research Center; Athens University of Economics Business(阿基米德雅典娜研究中心; 雅典娜研究中心语言与语音处理研究所; 雅典经济与商业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出希腊成文法检索基准GreekBarRetrieval,通过实验发现基于大型语言模型的查询重表述可提升BM25检索效果,使其在多项指标上优于其他检索方法。

AI 中文摘要

成文法检索对于基于引用的法律问答是必要的,但在希腊语领域仍未得到充分探索。我们推出了GreekBarRetrieval,这是一个源自并补充了GreekBarBench的公开检索基准,GreekBarBench此前未包含检索任务。该新基准包含283个律师资格考试问题,每个问题都附带其所指案件的事实,以及6308篇待检索的候选成文法条文。问题和事实以日常语言表述,但需要映射到成文法的正式术语及其抽象法律概念。另一个复杂因素是,并非所有案件事实都与案件的每个问题相关。我们对三种BM25变体和九种密集检索器进行了实验,发现普通密集检索在Recall@100指标上远优于普通稀疏检索。然而,基于大型语言模型(LLM)的查询重表述有助于缩小BM25与密集检索之间的差距,同时也能提升密集检索的效果。通过我们引入的十轮类ReAct的LLM重表述循环,BM25在Recall@100上进一步提升,并在所有测试的检索器中获得最佳的nDCG和MAP分数。查询重表述的效果也优于伪相关反馈、稀疏-密集融合以及英语翻译。

英文摘要

Statutory retrieval is necessary for citation-grounded legal question answering, but remains underexplored for Greek. We introduce GreekBarRetrieval, a public retrieval benchmark derived from, and complementing GreekBarBench, which did not include retrieval. The new benchmark comprises 283 bar-exam questions, each accompanied by the facts of the case it refers to, and 6,308 candidate statutory articles to retrieve from. Questions and facts are stated in everyday language, but need to be mapped to the formal terminology of statutes and their abstract legal concepts. A further complication is that not all of the case facts are relevant to each question of a case. Experimenting with three BM25 variants and nine dense retrievers, we find that vanilla dense retrieval far outperforms vanilla sparse retrieval in Recall@100. However, LLM-based query reformulation helps BM25 close that gap, while also improving dense retrieval. With a ten-round ReAct-like LLM reformulation loop that we introduce, BM25 improves further in Recall@100 and obtains the best nDCG and MAP scores of all tested retrievers. Query reformulation also outperforms pseudo-relevance feedback, sparse-dense fusion, and English translation.

CommentsAccepted at NLLP 2026. OpenReview: https://openreview.net/forum?id=LNK2RetzG8

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑