arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当故障传播时:智能体检索增强生成中的因果故障归因

When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation

Lauren Pothuru

arXiv 2608.20627首次发表:更新:

发表机构

Anote(阿诺特)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AgenticRAG-FP基准,针对智能体RAG的因果故障归因开展实验,发现基于覆盖率的诊断在第1轮表现较好,第2、3轮较差,明确将传播深度作为相关评估维度。

AI 中文摘要

智能体检索增强生成(RAG)会在多轮交互中交替执行检索、推理和答案生成。第1轮的检索错误可能仅在第3轮表现为错误答案,而后续的检索也可能修复该轨迹。本文提出AgenticRAG-FP,这是一个用于智能体RAG中因果故障归因的干预基准。该基准会在指定轮次注入经验证的故障,重新执行下游轨迹,并基于已知干预对诊断器进行评估。其核心问题是,在后缀发生变化后,事后追踪是否仍能识别出注入故障的轮次。在对80个三轮MuSiQue问题完成的严格密集Claude Haiku 4.5扫描中,基于覆盖率的诊断在第1轮为0.91,在第2、3轮为0.00(n=43、36、21条失败轨迹)。一项规模较小的内容损坏研究会在主题完整的证据中改变承载答案或桥梁事实,在深度2(过滤后剩余18个失败案例),基于覆盖率的诊断为0.00,而冻结轮次反事实探针在探索性合并比较中为0.67。深度3的内容估计仅具描述性,因其仅包含3个失败案例。这些结果将传播深度明确作为诊断智能体RAG故障的评估维度,同时区分事后信号丢失的广泛证据与小样本方法比较。

英文摘要

Agentic retrieval-augmented generation (RAG) interleaves retrieval, reasoning, and answer generation across multiple hops. A retrieval error at hop 1 can surface only as a wrong answer at hop 3, while later retrieval can also repair the trajectory. This paper introduces AgenticRAG-FP, an interventional benchmark for causal failure attribution in agentic RAG. The benchmark injects a certified fault at a specified hop, re-executes the downstream trajectory, and evaluates diagnosers against the known intervention. Its central question is whether a post-hoc trace still identifies the injected hop after the suffix changes. In the completed strict dense Claude Haiku 4.5 sweep on 80 three-hop MuSiQue questions, coverage-based diagnosis is 0.91 at hop 1 and 0.00 at hops 2 and 3 (n=43,36,21 failed trajectories). A smaller content-corruption study changes an answer-bearing or bridge fact in topically intact evidence. At depth 2, where 18 failed cases remain after filtering, coverage-based diagnosis is 0.00 and a frozen-hop counterfactual probe is 0.67 in an exploratory pooled comparison. Depth-3 content estimates are descriptive only because they contain three failed cases. These results make propagation depth an explicit evaluation axis for diagnosing agentic RAG failures while distinguishing broad evidence of post-hoc signal loss from small-sample method comparisons.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑