发表机构
Duke University; University of Pennsylvania(杜克大学; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SearchAtlas将搜索轨迹转为证据查询图,以86.0%边F1解析,揭示智能体搜索过程缺陷,其诊断分数比LLM评判更关联错误答案,提升可解释性。
AI 中文摘要
LLM搜索智能体通常仅以最终答案的准确性进行评估,忽略了过程。分析搜索策略需要理解如何检索可信证据以满足问题约束。这些有价值的信息深埋在冗长且难以解析的原始搜索轨迹中。我们引入了SearchAtlas,一个将搜索轨迹转换为结构化图的框架,其边表示证据如何在推理轨迹中传播,从检索证据的查询到最终答案。我们的自动化解析流水线在人工标注图上实现了86.0%的平均边F1分数,并在多次运行中保持一致。我们分析了三个基准上的五个搜索智能体,揭示了搜索规模和证据聚合的系统性差异。SearchAtlas暴露了碎片化的答案支持、未到达答案的问题约束以及未经验证的参数化知识进入响应的问题。这些过程失败与错误答案密切相关,其关联程度甚至高于仅基于原始轨迹或有序查询列表的LLM评判器,表明构建的图提供了有用的可解释性。此外,对过程诊断分数与最终答案正确性不一致的案例进行审计显示,这些分数捕获了无法简化为答案准确性的信息。
英文摘要
LLM search agents are often evaluated on final-answer accuracy, overlooking the process. Analyzing a search strategy requires understanding how credible evidence is retrieved to address question constraints. This valuable information is buried in raw search trajectories that are long and difficult to parse. We introduce SearchAtlas, a framework that converts search trajectories into structured graphs whose edges represent how evidence is propagated across the reasoning trace, from the query that retrieves it to the final answer. Our automated parsing pipeline achieves a mean edge F1 of 86.0% against human-annotated graphs and remains consistent across repeated runs. We analyze five search agents on three benchmarks, revealing systematic differences in search scale and evidence aggregation. SearchAtlas exposes fragmented answer support, question constraints that do not reach the answer, and unverified parametric knowledge entering the response. These process failures are strongly associated with incorrect answers, even more so than an LLM judge given either the raw trajectory or the ordered query list, suggesting that the constructed graphs provide useful interpretability. Moreover, an audit of cases in which process-diagnostic scores disagree with final-answer correctness shows that they capture information not reducible to answer accuracy.
CommentsAccepted to Findings of EMNLP 2026. 30 pages, 9 figures, including appendices