SearchAuditor:长程搜索智能体失败的审计与归因
SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
浏览论文内容
中文总结 AI 辅助
本研究针对长程搜索智能体的失败诊断难题,构建SearchAuditBench基准,提出多视角审计框架SearchAuditor,其端到端通过率达32.3%,优于基线模型,可助力智能体从错误中恢复。
中文摘要 AI 辅助
深度搜索智能体通过长程网络交互解决具有挑战性的问题,这一过程既复杂又脆弱:微小的推理错误可能会在漫长且充满噪声的轨迹中传播,最终形成流畅但错误的答案。诊断此类失败十分困难,需要人工检查极长的执行轨迹,这可能超出人类的能力范围。因此,我们引入SearchAuditBench,这是一个评估大语言模型(LLM)审计器能否定位、归因和修复这些失败的基准,旨在减轻人类负担。SearchAuditBench包含1243条失败轨迹,平均每条有73.1条消息和65.1K个token,这些轨迹是从8个开放权重模型在5个深度搜索基准上收集的,每条轨迹都由专家标注了关键错误步骤、搜索特定根本原因以及带有评分规则的参考修复方案。我们进一步提出SearchAuditor,这是一个多视角审计框架,通过基于证据的裁决有效定位、归因和修复搜索智能体的失败。实验结果显示,即使是最强的基线模型,在GPT-5.5等前沿模型的支持下,其端到端通过率也仅为26.6%。相比之下,我们的SearchAuditor在不同前沿模型上始终优于所有基线模型,实现了32.3%的端到端通过率,且其修复方案能让智能体更好地从错误中恢复,从而重新运行失败的任务。
英文摘要
Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answers. Diagnosing such failures is difficult, requiring the manual inspection of extremely long execution traces, which could be beyond human capacity. We therefore introduce SearchAuditBench, a benchmark that evaluates whether LLM auditors can localize, attribute, and repair these failures, thereby reducing the human burden. SearchAuditBench comprises 1,243 failed trajectories, averaging 73.1 messages and 65.1K tokens, collected from eight open-weight models on five deep-search benchmarks, each expert-annotated with the critical error step, a search-specific root cause, and a reference repair with grading rubrics. We further propose SearchAuditor, a multi-perspective auditing framework that effectively localizes, attributes, and repairs search-agent failures through evidence-grounded adjudication. Experimental results show that even the strongest baseline, when powered by a frontier model like GPT-5.5, attains only a 26.6% end-to-end pass rate. In contrast, our SearchAuditor consistently outperforms all baselines across different frontier models, achieving an end-to-end pass rate of 32.3%, and resuming failed runs with its repairs enables agents to better recover from errors.