问题的开局:第一步在智能体深度搜索中的重要性
Question's Gambit: The First Move Matters in Agentic Deep Search
- University of Toronto(多伦多大学)
- Mila – Quebec AI Institute(米拉-魁北克人工智能研究所)
- University of Waterloo(滑铁卢大学)
- University of California, Berkeley(加州大学伯克利分校)
- McGill University(麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出Question's Gambit首步检索模块,通过分解线索、重构搜索并重排序,提升深度搜索智能体在BrowseComp-Plus上的答案准确率至90.5%,验证了首步质量的关键作用。
AI中文摘要:
深度研究智能体通过搜索、阅读和推理的迭代循环来回答复杂问题。近期在诸如BrowseComp-Plus等推理密集型基准上的工作表明,配置良好的词汇检索能够提供高质量的证据,但智能体仍可能无法将携带证据的文档与黄金文档关联起来。我们确定深度研究智能体的首次检索动作是该场景中的一个重要设计决策。我们引入了Question's Gambit,一个首步检索模块,它将问题分解为一组线索,将其重构为互补的搜索,整合检索结果,并在智能体开始其迭代搜索与推理过程之前对候选池进行重排序。这产生了一个开局上下文,旨在同时支持线索聚合和最终答案验证。我们进一步在MultiHop-RAG上评估,以测试这些优势是否能超越BrowseComp-Plus,迁移到更传统的多跳问题结构。在BrowseComp-Plus上的实验表明,Question's Gambit相较于强基线提高了检索召回率和下游智能体准确率,在使用gpt-5.5时,将答案准确率从83.1%提升至90.5%,超过了最强报告的智能体基线Pi-Serini。我们的结果证实,有效的智能体深度搜索不仅依赖于循环内可用的工具,还取决于首步的质量。我们在以下网址公开了我们的实现:此https URL。
英文摘要:
Deep research agents answer complex questions through iterative loops of searching, reading, and reasoning. Recent work on reasoning-intensive benchmarks such as BrowseComp-Plus shows that well-configured lexical retrieval can surface high-quality evidence, yet agents may still fail to connect documents carrying evidence to the gold documents. We identify a deep research agent's first retrieval move as an important design decision for this setting. We introduce Question's Gambit, a first-move retrieval module that decomposes the question into a set of clues, reformulates them into complementary searches, consolidates the retrieved results, and reranks the candidate pool before the agent begins its iterative search-and-reasoning process. This produces an opening context designed to support both clue aggregation and final-answer verification. We further evaluate on MultiHop-RAG to test whether these benefits transfer beyond BrowseComp-Plus to a more conventional multi-hop question structure. Experiments on BrowseComp-Plus show that Question's Gambit improves retrieval recall and downstream agent accuracy over strong baselines, improving answer accuracy from 83.1% to 90.5% with gpt-5.5 over Pi-Serini, the strongest reported agentic baseline. Our results confirm that effective agentic deep research depends not only on the tools available inside the loop, but also on the quality of the first move. We published our implementation publicly at https://github.com/radinhamidi/Question-s-Gambit.