深度压力测试:对深度搜索智能体进行压力测试
DeepStress: Stress-Testing Deep Search Agents
浏览论文内容
中文总结 AI 辅助
研究搜索智能体对低质量证据的鲁棒性,提出深度压力测试框架DeepStress,通过可控合成环境替换检索模块控制挑战性证据频率,在HotpotQA和BrowseCompPlus上测试,发现智能体处理不可靠信息能力有差异并提出新指标。
中文摘要 AI 辅助
虽然搜索智能体在多步问答中展现出强大能力,但它们对低质量证据的鲁棒性仍未得到充分探索。这种现象在现实基准测试中很少出现,但在实际应用中可能导致严重失败。因此,本研究提出了深度压力测试(DeepStress),这是一个压力测试框架,通过用可控的合成环境替换搜索智能体的检索模块来控制挑战性证据的频率。我们用这个框架控制三个可影响文档可靠性的维度:可信度、相关性和事实性。在HotpotQA和BrowseCompPlus上对多个搜索智能体进行测试,结果表明智能体在处理不可靠信息的能力上存在显著差异,并提出了新的指标来更好地记录系统结果以及冲突的参数知识和检索到的知识之间的相互作用。
英文摘要
While search agents demonstrate impressive capabilities in multi-step question answering, their robustness to poor-quality evidence remains under-explored. This phenomenon occurs rarely in realistic benchmarks but can lead to dramatic failure in real life applications. Therefore in this study we propose DeepStress, a stress testing framework that controls the frequency of challenging evidence by replacing the retrieval module of search agents with a controlled synthetic environment. We use this framework to control three dimensions that can affect document reliability: trustworthiness, relevance, and factuality. Testing several search agents on HotpotQA and BrowseCompPlus, we demonstrate that agents exhibit substantial differences in their ability to handle unreliable information and propose new metrics that better document systems outcomes as well as the interactions between conflicting parametric and retrieved knowledge.