arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30929cs.CL

用于波兰成文法的带注释代理检索

Annotated Surrogate Retrieval for Polish Statutory Law

Orkun Yiğit Cengiz

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出基于文档代理的波兰成文法检索方法(ASCR、ASCR-H、DTF),在法律考试问题数据集上验证,ASCR-H排名1准确率达72.3%,DTF在排名20时领先且成本更低,同时公开了相关基准数据。

中文摘要 AI 辅助

我们提出了一系列基于文档代理的波兰成文法检索方法:在索引时将语言模型注释附加到成文法条款上。三种设计在成本-质量权衡曲线上占据不同位置:ASCR是带重排序的代理级联;ASCR-H将密集列表融合到该级联中;DTF则用三个词汇检索器、密集检索器、加权互反秩融合以及确定性重评分先验替代了两个语言模型阶段,在生成前不使用任何模型调用。我们针对14个词汇、密集、融合和消融基线以及4个对照,在来自2024和2025年波兰律师及法律顾问入职考试的300个问题(其中264个问题的参考条款在语料库中)上对这三种方法进行了评估,语料库包含来自1133部法案的82508个条款。在配对McNemar检验中,除自身一个消融配置外,ASCR-H将参考条款排在首位的频率显著高于其他所有非神谕配置(20次比较中有18次在p<0.005水平下对其有利),达到72.3%,而BM25为61.7%,密集检索为52.3%。该优势集中在排名前列,随排名深度增加而消失:在排名1和5时具有显著性,排名10时消失,排名20时DTF在点估计上领先(86.0%对84.5%),延迟仅为其1/9,成本不到其一半。消融实验显示,仅重排序阶段就贡献了27.6个百分点的排名1准确率。我们还发现,排名优势未延伸至引用准确率,其中DTF达到神谕上限,此外还报告了关于词形还原、伪相关反馈和查询重写的三个负面结果。代理注释覆盖了27.0%的语料库,但包含基准中的所有参考条款,我们披露并讨论了这种不对称性。基准、每个问题的输出以及配对显著性检验均已公开。

英文摘要

We present a family of retrieval methods for Polish statutory law built on document surrogates: language-model annotations attached to statutory articles at index time. Three designs occupy different points on the cost-quality frontier. ASCR is a surrogate cascade with reranking; ASCR-H fuses a dense list into that cascade; and DTF replaces both language-model stages with three lexical and dense retrievers, weighted reciprocal rank fusion, and a deterministic re-scoring prior, using no model call before generation. We evaluate all three against fourteen lexical, dense, fused and ablated baselines plus four controls, on 300 questions from the 2024 and 2025 Polish bar and legal counsel entrance examinations (264 with their reference article in the corpus), over 82,508 articles from 1,133 acts. On paired McNemar tests, ASCR-H places the reference provision at rank one significantly more often than every other non-oracle configuration except one of its own ablations (eighteen of twenty comparisons significant in its favour at p < 0.005), reaching 72.3% against 61.7% for BM25 and 52.3% for dense retrieval. The advantage is concentrated at the head and does not survive depth: it is significant at cutoffs of one and five, disappears by ten, and by twenty DTF leads on point estimate (86.0% versus 84.5%) at one ninth the latency and less than half the cost. Ablation attributes 27.6 points of rank-one accuracy to the reranking stage alone. We further report that the ranking advantage does not extend to citation accuracy, where DTF matches the oracle ceiling, and three negative results on lemmatisation, pseudo-relevance feedback and query rewriting. Surrogate annotation covers 27.0% of the corpus but every reference provision in the benchmark, an asymmetry we disclose and discuss. Benchmark, per-question outputs and paired significance tests are publicly available.

补充信息

↑