证据-推理重建:当证据被回忆但推理出错时
Evidence-Inference Reconstruction: When The Evidence Is Recalled But The Reasoning Goes Wrong
- University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
提出证据-推理重建(EIR)方法,通过结构化状态引导单次检索并累积证据,在最终调用中生成答案,显著提升多跳问答准确率并大幅减少模型调用次数。
中文摘要 AI 辅助
现代多跳LLM智能体配备了内置机制来检测中间推理步骤中的错误。此类错误会触发智能体的纠正行动,这些行动大多遵循重试步骤或推理轨迹的范式。这些重试不仅代价高昂,我们在本文中还表明它们可能是不必要的。为此,我们引入了证据-推理重建(EIR),它使用结构化状态来引导一次检索轨迹,在此过程中累积源证据。我们表明,只要相关证据已被收集,即使由于错误的中间推理步骤而混入了错误证据,EIR也能在单次最终模型调用中生成正确答案。在一项评估中,使用Haiku 4.5和GPT-4.1 Mini,我们在HotpotQA、2WikiMultiHopQA和MuSiQue的匹配的1,000个问题子集上评估了EIR,表明EIR在Answer F1(模型答案与正确答案的重叠)上比基线提高了8.3–32.8个百分点,比Agentic SSR提高了10.6–29.1个百分点,比Reflexion提高了1.1–15.9个百分点。此外,我们表明EIR每个问题平均总模型调用次数为4.85次,而Agentic SSR为35.29次,Reflexion为12.41次。这些结果共同证实了EIR的核心前提:将证据检索与最终答案模型调用分离,可以在利用显著更少计算量的同时提高答案准确性。
英文摘要
Modern multi-hop LLM agents are equipped with built-in mechanisms to detect errors in intermediate reasoning steps. Such errors trigger corrective actions from these agents, which mostly follow the paradigm of retrying the steps or the reasoning trajectories. Not only are these retries expensive, we present in this paper that they are also potentially unnecessary. To this end, we introduce Evidence-Inference Reconstruction (EIR), which uses structured state to guide one retrieval trajectory, accumulating source evidence in the process. We show that as long as the relevant evidence has been collected, EIR is capable of generating the correct answer in a single final model call even if erroneous evidence has been mixed in due to incorrect intermediate reasoning steps. In one evaluation, using Haiku 4.5 and GPT-4.1 Mini, we evaluate EIR on matched 1,000-question subsets of HotpotQA, 2WikiMultiHopQA, and MuSiQue, showing that EIR improves Answer F1, the overlap between the model's and the correct answer, over the baseline by 8.3--32.8 points, Agentic SSR by 10.6--29.1 points, and Reflexion by 1.1--15.9 points. Additionally, we show that EIR averages 4.85 total model calls per question, compared with 35.29 for Agentic SSR and 12.41 for Reflexion. Together, these results corroborate EIR's central premise: separating evidence retrieval from the final answer model call can improve answer accuracy while utilizing substantially less computation.