发表机构
Bar-Ilan University; UNC Chapel Hill; University of Texas at Austin(巴伊兰大学; 北卡罗来纳大学教堂山分校; 德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对智能体深度研究系统的引用召回率低问题,提出定位错误来源的评估方法与四类错误分类法,应用于三个开源系统发现协调器为主要错误来源,通过简单干预提升5%引用召回率且不降低质量。
AI 中文摘要
深度研究(Deep Research, DR)系统通过协调多个智能体从网络搜索并合成信息,生成带引用的长篇报告。引用是评估这些报告忠实度的主要机制,但当前DR系统的引用召回率较低。此外,提升引用召回率颇具挑战性,因为DR系统是复杂的多智能体架构,信息会像传话游戏一样在智能体间传递,内容和引用都可能在此过程中出错。我们提出一种评估方法,通过针对各智能体调用相对于自身输入的忠实度和可验证性进行局部测试,精准定位是哪个智能体引入了错误。此外,我们提出一种四类分类法,对发现的错误进行分类:幻觉、依赖未引用输入、输出未引用、引用不足。将我们的方法应用于三个排名靠前的开源DR系统,我们获得了可操作的诊断结果。除了汇总单一文档的智能体外,几乎每个智能体都犯了大量错误。我们发现,不同智能体的主要错误类型存在系统性差异,其中协调器的错误大多与引用相关。我们发现,AI-Q中84.7%的最终报告错误源于协调器,其中约31%是幻觉,其余为引用错误。基于这些见解,我们证明两种简单干预措施可将引用召回率提升5%,且不会降低输出质量。
英文摘要
Deep research (DR) systems produce long-form cited reports by orchestrating multiple agents that search and synthesize information from the web. Citations are the primary mechanism for evaluating the faithfulness of these reports, yet current DR systems exhibit poor citation recall. Moreover, improving citation recall is challenging because DR systems are complex multi-agent architectures where information passes through agents like a telephone game, and both content and citations can get corrupted along the way. We propose an evaluation method that pinpoints which agent introduced each error by locally testing agent invocations for faithfulness and verifiability relative to their own inputs. Furthermore, we propose a four-type taxonomy to categorize the discovered errors: hallucination, uncited input reliance, uncited output, or insufficient citations. Applying our method to three top-ranked open-source DR systems, we obtain actionable diagnostics. Almost every agent makes a lot of mistakes with the exception being those that summarize a single document. We find that the dominant error type varies systematically across agents, where the orchestrator mistakes are mostly citation-related. We find that 84.7% of final-report errors in AI-Q originate at the orchestrator, roughly 31% of them hallucinations and the rest citation mistakes. Guided by these insights, we demonstrate that two simple interventions raise citation recall by 5% without degrading output quality.
CommentsAccepted to EMNLP 2026 (Main Conference). Code: https://github.com/eranhirs/who-is-the-agent-to-blame