arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31342cs.CL

过期文档投毒:过时检索覆盖模型正确答案

Stale-Document Poisoning: When Outdated Retrieval Overrides Correct Model Answers

Md Shamim Ahmed, Lukas Galke Poech, Richard Röttger

首次发表
浏览论文内容

中文总结 AI 辅助

本研究识别RAG中的过期文档投毒问题,构建317个知识反转基准,证明过时检索可显著翻转模型答案,并提出依赖时序元数据的重排序器以缓解该问题。

中文摘要 AI 辅助

检索增强生成(RAG)常被用于通过提供外部证据来应对知识过时的问题。但检索仅在证据仍然有效时才有帮助。我们识别出一种时间对齐失败——过期文档投毒,即过时证据使模型出错,尽管在没有检索的情况下模型能正确回答。我们构建了一个包含317个经核实的知识反转的基准,涵盖医学、法律、软件和平台政策领域,基于注明日期的官方来源。在12个模型中,近期医学反转比长期存在的反转更难处理。更重要的是,过时检索在无指示信任文档的情况下使30%的Llama和37%的Qwen答案发生反转;明确的跟随指示将这些比例提升至66%和75%。在四个开放模型和四个领域中,投毒率从17%到91%不等,而匹配的最新证据在97-100%的试验中被遵循。为隔离时间适用性,我们在50个反转中保持历史证据不变,仅改变评估日期。一个清晰的模式浮现:仅日期产生适度的适应,但当模型被明确告知旧证据何时停止适用时,较大的模型几乎完美地切换到适当答案。因果干预确认这一有效性信息直接塑造最终决策。相同的内部组件也支持更广泛的比较任务,表明时间适用性可以调用用于其他比较的通用推理机制。最后,一个固定的、感知时效性的混合重排序器在日期准确时将投毒率降低4.6-10.0个百分点,其收益依赖于可靠的时序元数据。因此,可靠的RAG需要选择性信任:模型不仅必须确定检索证据说了什么,还必须确定它是否仍然适用。

英文摘要

Retrieval-augmented generation (RAG) is often used to address outdated knowledge by providing external evidence. But retrieval helps only when that evidence is still valid. We identify a temporal alignment failure, stale-document poisoning, in which outdated evidence makes a model wrong despite answering correctly without retrieval. We construct a benchmark of 317 verified knowledge reversals across medicine, law, software, and platform policy, grounded in dated official sources. Across 12 models, recent medical reversals are harder than long-established ones. More importantly, outdated retrieval flips 30% of Llama and 37% of Qwen answers even without instructions to trust the document; explicit follow instructions raise these rates to 66% and 75%. Across four open models and four domains, poisoning ranges from 17-91%, while matched up-to-date evidence is followed in 97-100% of trials. To isolate temporal applicability, we keep the historical evidence unchanged across 50 reversals and vary only the evaluation date. A clear pattern emerges: dates alone produce only modest adaptation, but when models are explicitly told when the old evidence stops applying, the larger models switch to the appropriate answer almost perfectly. Causal interventions confirm that this validity information directly shapes the final decision. The same internal components also support broader comparison tasks, suggesting that temporal applicability can recruit a general reasoning mechanism used for other comparisons. Finally, a fixed recency-aware hybrid re-ranker reduces poisoning by 4.6-10.0 points when dates are accurate, with gains that depend on reliable temporal metadata. Reliable RAG therefore requires selective trust: models must determine not only what retrieved evidence says, but whether it still applies.

发表机构

  • University of Southern Denmark(南丹麦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑