arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

法律检索增强生成(RAG)中的时间错位:面向法国税法的版本化语料库基准

Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law

Rose Cymbler, Daniel Guez, Laurent Fabre

arXiv 2608.09393首次发表:更新:

发表机构

Talia; Databricks(塔利亚; Databricks公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对法律RAG的时间错位问题,构建法国税法版本化语料库基准FiscalQA Pro,发现现有模型时间推理能力差,端到端多版本检索器性能优异,还发布了相关数据集与代码。

AI 中文摘要

我们识别并量化了时间错位问题:当适用版本为更早或未来版本时,系统会系统性地检索并引用现行有效的法律条文版本。标准法律RAG将语料库视为静态的,我们认为法律问答是一个需按时间索引的检索问题。我们推出了FiscalQA Pro,它将包含32436条法国税法条文版本(时间跨度93年,1938年至2031年)的版本化语料库,与一个全模型难度的时间推理任务轨道相结合:该轨道包含33条CGI条目中的209个经评分、专家审核的问题(其中221个已发布,12个被标记为超出可回答范围)。在选择阶段,所有被评估的模型在4次抽样中均未以闭卷形式恢复其日期适用的答案,且现行文本仅对1个已评分问题包含黄金标准值。答案通过原子真实值“ nuggets”(正则表达式和带容差的数值)进行确定性评分,绝不使用大模型作为评判者:因为大模型评判者会继承其要评分的时间偏差。在11个模型(5个前沿闭源API系统、作为替代参赛项的Gemini 2.5 Pro,以及5个开源权重模型)中,参数知识的平均严格准确率为3.0%,对静态现行版本语料库的RAG为2.7%。静态RAG在0%的时间内检索到日期适用的版本,自信地引用了真实但不适用的版本。我们的端到端检索器在多版本索引上运行,无需神谕,达到了98.3%的平均严格准确率;神谕条文消融实验达到99.1%,将剩余差距定位在第一阶段召回,而非版本选择。我们还发布了一个包含69208个引用链接的版本化判例数据集,以及该语料库、基准、模型响应和管道代码。

英文摘要

We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one. Standard legal RAG treats the corpus as static; we argue legal question answering is a temporally-indexed retrieval problem. We introduce FiscalQA Pro, pairing a versioned corpus of 32,436 article-versions of the French tax code (93 years, 1938-2031) with an all-model-hard temporal-reasoning track: 209 scored, expert-reviewed questions across 33 CGI articles (221 released; twelve flagged out of the answerable scope). At selection time, no evaluated model recovered its date-applicable answer closed-book in any of four sampling draws, and the currently in-force text lacks the gold value for all but one of the scored questions. Answers are scored deterministically via atomic ground-truth "nuggets" (regex and numeric-with-tolerance), never LLM-as-judge: an LLM judge would inherit the temporal bias it is meant to score. Across eleven models (five frontier closed-API systems plus Gemini 2.5 Pro as a substitute entry, and five open-weight), parametric knowledge yields 3.0% mean strict accuracy and RAG over a static current-version corpus 2.7%. Static RAG retrieves the date-applicable version 0% of the time, confidently citing a real but inapplicable version. Our end-to-end retriever over a multi-version index, with no oracle, reaches 98.3% mean strict; an oracle-article ablation reaches 99.1%, locating the residual gap in first-stage recall, not version selection. We additionally release a version-aware jurisprudence dataset of 69,208 citation links, together with the corpus, benchmark, model responses, and pipeline code.

Comments13 pages, 1 figure, 4 tables. Accepted at the ICML 2026 Workshop on AI for Law (AI4Law), Seoul. Code and data: https://github.com/rosecymbler/fiscal-fr-bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑