arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在推理失效之前:智能体检索增强生成(RAG)中的证据前程序失效

Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG

Daeyoung Roh, Donghee Han

arXiv 2608.02011首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对智能体RAG系统在推理前的程序失效,提出将错误答案分为两类失效,评估Read-Gate机制,发现强制阅读可提升准确率,且证据收集应作为独立轨迹控制问题评估。

AI 中文摘要

智能体检索增强生成(RAG)系统可能在以证据为条件的推理测试前就出现失效:智能体可能检索到候选片段,但在未检查它们的情况下就得出最终答案。我们将这种失效模式视为智能体轨迹的程序属性,利用保存的工具调用轨迹、检索到的证据、已阅读的段落和最终答案,将错误答案分解为证据前纪律失效和金标准阅读后失效。在HotpotQA、2WikiMultiHopQA和MuSiQue上的12000条配对轨迹中,两种失效类型基本不重叠:在正则表达式和spaCy实体提取器下,两者同时触发的比例在11.2%至13.1%之间。我们随后评估了Read-Gate,这是一种要求智能体在搜索后、最终得出结论前必须阅读的最小运行时不变量。强制阅读在原本会跳过阅读的轨迹上将LLM准确率(LLM-Acc)提高了14.9至19.9个百分点,在完整的最小推理单元上提高了3.2至9.4个百分点。额外的诊断显示,更大的隐藏思考预算并不一定会增加证据检查。总体而言,这些结果表明,证据收集应作为轨迹级别的控制问题进行评估,与答案侧推理分开处理。

英文摘要

Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them. We study this failure mode as a procedural property of the agent trajectory, decomposing wrong answers into pre-evidence discipline failures and post-gold-read failures using saved tool-call traces, retrieved evidence, read passages, and final answers. Across 12,000 paired trajectories on HotpotQA, 2WikiMultiHopQA, and MuSiQue, the two failure types are largely non-redundant: the both-trigger rate is in [11.2%, 13.1%] across regex and spaCy entity extractors. We then evaluate Read-Gate, a minimal runtime invariant requiring an agent to read after search and before finalization. Forced reading improves LLM-Acc by 14.9-19.9 points on trajectories that would otherwise skip reading and by 3.2-9.4 points on full minimal-reasoning cells. Additional diagnostics show that larger hidden thinking budgets do not necessarily increase evidence inspection. Together, these results indicate that evidence-gathering should be evaluated as a trajectory-level control problem, separately from answer-side reasoning.

Comments22 pages, 7 figures. Code: https://github.com/Noverse0/before-reasoning-fails

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑