arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22856cs.IRcs.CL

同一智能体,不同答案:对检索增强问答中语料库诱导的答案波动的感知重复审计

Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA

Jingjie Ning, Xueqi Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对检索增强问答系统索引扩展后出现的答案波动现象,提出Snapshot Compatibility Audit方法,经Natural Questions、TriviaQA等数据集验证,发现存在无对应准确率变化的答案兼容性失效问题,建议检索增强发布需审计兼容性。

中文摘要 AI 辅助

检索增强问答(QA)系统在其请求的模型标识符、提示、检索策略、证据深度、渲染和公开的生成控制均保持固定的情况下,即使在索引扩展后也可能返回不同的答案。当收益和损失相互抵消时,总体准确率可能会隐藏这些变化,而普通的生成变异性会使单次比较夸大更新效果。我们将这种隐藏现象称为准确率盲答案波动,并引入“快照兼容性审计”(Snapshot Compatibility Audit),该方法通过将跨快照分歧减去同一快照内的重复分歧来估算额外答案波动。我们通过将一个冻结的FineWeb前缀从1个分片扩展到7个分片来实例化该方法。在一项预先注册的400题Natural Questions研究中,标准化精确匹配和盲语义额外波动分别为6.44和10.25个百分点,而精确匹配准确率仅变化了-1.50个百分点。一项事后分析发现,400题中有40题存在重复稳定的语义翻转。另一项预先注册的200题TriviaQA研究产生了更小、方向一致的额外波动,而精确匹配准确率则向相反方向变化。一项采用第二个DeepSeek生成器和服务配置的无结果事后100题子集复制,发现语义额外波动为8.75个百分点,而精确匹配上升了3.00个百分点。因此,答案级别的兼容性可能在没有明显或一致的效用变化的情况下失效。检索增强版本应在效用之外审计兼容性。

英文摘要

A retrieval-augmented QA system can return different answers after an index expansion even when its requested model identifier, prompt, retrieval policy, evidence depth, rendering, and exposed generation controls are held fixed. Aggregate accuracy may hide these changes when gains and losses cancel, while ordinary generation variability makes one-shot comparisons overstate update effects. We call the hidden phenomenon accuracy-blind answer churn and introduce the \emph{Snapshot Compatibility Audit}, which estimates excess answer churn by subtracting same-snapshot repeat disagreement from cross-snapshot disagreement. We instantiate it by expanding one frozen FineWeb prefix from one to seven shards. In a preregistered 400-question Natural Questions study, normalized-exact and blinded-semantic excess churn are 6.44 and 10.25 percentage points while exact-match accuracy changes by only $-1.50$ points. A post-hoc analysis finds repeat-stable semantic flips on 40/400 questions. A separately preregistered 200-question TriviaQA study yields smaller, directionally consistent excess churn while exact-match accuracy moves in the opposite direction. An outcome-blind post-hoc 100-question subset replication with a second DeepSeek generator and serving configuration finds 8.75 pp of semantic excess churn even as exact match rises by 3.00 percentage points. Answer-level compatibility can therefore fail without a conspicuous or consistently directed utility shift. Retrieval-augmented releases should audit compatibility alongside utility.

发表机构

  • School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

↑