arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

如何修复此RAG失败?通过配对证据干预审计反事实响应

What Would Fix This RAG Failure? Auditing Counterfactual Response with Paired Evidence Interventions

Wenzhang Du

arXiv 2608.08944首次发表:更新:

AI 中文总结

本研究提出Pair-ID离线审计方法,通过配对证据干预测量RAG失败的反事实响应,实验发现证据敏感性存在且部分可预测、依赖阅读器,支持离线响应审计框架。

AI 中文摘要

检索增强生成(RAG)的失败回答可能与多种未见过的证据修复响应一致。我们提出Pair-ID,一种离线审计方法,其保持同一查询、检索状态和阅读器不变,然后交叉两种操作——添加缺失的支持证据和删除已验证的非支持证据——以测量同一失败的反事实响应向量。在19981个基准查询上的完整漏斗筛选出11105个符合条件的通义千问(Qwen)失败案例,从中采用前瞻性修复的SHA-256排序选择1200个案例生成任意采样响应。在1190个重新生成的有效失败案例中,添加支持证据修复了600个JOINT案例中的197个(修复率0.328,95%置信区间[0.292, 0.367]),删除操作修复了1190个案例中的162个(修复率0.136,95%置信区间[0.117, 0.155]);长度和位置匹配的假操作保留了0.223和0.101的语义对比。原始视图对单个响应单元具有部分预测信号(宏AUROC为0.678;Brier得分为0.152,而边际基线为0.160),但精确向量准确率0.637未超过0.646的多数向量基线,向量宏F1为0.170。在四个阅读器中,两种边际敏感性均重复出现,而合并的精确向量一致性为0.675-0.765,仅JOINT的一致性降至0.538-0.691。这些结果表明,在哈希筛选的符合条件的失败样本中,证据敏感性以有意义的比例存在,仅能从观察到的失败中部分预测,且取决于阅读器。该证据支持范围限定的离线响应审计框架,而非信息论的不可能结果、独立于阅读器的分类法或运行时修复策略。

英文摘要

A failed retrieval-augmented generation (RAG) answer can be consistent with several unseen responses to evidence repair. We introduce Pair-ID, an offline audit that holds one query, retrieval state, and reader constant, then crosses two operations, adding missing support and deleting verified nonsupport, to measure a same-failure counterfactual response vector. A complete funnel over 19,981 benchmark queries identifies 11,105 eligible Qwen failures, from which a prospectively fixed SHA-256 ordering selects 1,200 before generating any sampled response. Among 1,190 regenerated-valid failures, support addition repairs 197/600 JOINT cases (0.328, 95% CI [0.292, 0.367]), and deletion repairs 162/1,190 cases (0.136, 95% CI [0.117, 0.155]); length- and position-matched shams retain semantic contrasts of 0.223 and 0.101. The original view carries partial predictive signal for individual response cells (macro AUROC 0.678; Brier 0.152 versus 0.160 for a marginal baseline), but exact-vector accuracy, 0.637, does not exceed the 0.646 majority-vector baseline, and vector macro-F1 is 0.170. Across four readers, both marginal sensitivities recur, while pooled exact-vector agreement is 0.675-0.765 and JOINT-only agreement falls to 0.538-0.691. These results show that evidence sensitivity occurs at meaningful rates in the hash-selected eligible-failure sample, is only partially predictable from the observed failure, and is conditional on the reader. The evidence supports a frame-scoped offline response audit, not an information-theoretic impossibility result, reader-independent taxonomy, or runtime repair policy.

Comments15 pages, 2 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑