ITER:面向智能体搜索的交互感知检索
ITER: Interaction-Aware Retrieval for Agentic Search
浏览论文内容
中文总结 AI 辅助
本文提出交互感知稠密检索器ITER,利用智能体轨迹学习信号训练,在六种智能体骨干上优于LRAT,跨智能体鲁棒性强于AgentIR,在两个基准上获显著性能提升。
中文摘要 AI 辅助
深度研究智能体通过迭代的搜索步骤序列回答复杂用户问题,智能体自主构造子查询以在每个阶段检索所需证据。然而,现有检索器训练通常仅依赖当前步骤的子查询及其对应搜索结果作为训练信号,导致从先前交互中积累的信息未被充分利用。我们提出ITER,一种利用智能体轨迹学习信号训练的智能体交互感知稠密检索器。ITER表示每个查询时不仅纳入当前子查询,还纳入主问题和先前子查询,并使用源自智能体交互的轨迹相对学习信号进行训练。在来自三个模型家族的六种智能体骨干上,ITER始终优于现有经智能体轨迹训练的稠密检索器LRAT,在InfoSeek-Eval上平均提升7.5%,在BrowseComp-Plus上平均提升13.5%。ITER还展现出比AgentIR更强的跨智能体鲁棒性,AgentIR是一种依赖外部LLM评判信号和智能体预搜索推理的深度研究检索器。消融实验进一步表明,主问题和先前子查询能提供最鲁棒的查询表示,而先前访问过的有用文档(在后续搜索中用作冗余负样本)能提供最强的轨迹相对监督。代码可在https URL获取。
英文摘要
Deep-research agents answer complex user questions through an iterative sequence of search steps, where the agent autonomously formulates sub-queries to retrieve the evidence needed at each stage. However, existing retriever training typically relies only on the sub-query and its corresponding search results at the current step as training signals, leaving the information accumulated from previous interactions largely underutilized. We introduce ITER, an agent interaction-aware dense retriever trained using agent trajectory learning signals. ITER represents each query by incorporating not only the current sub-query, but also the main question, the agent's pre-search reasoning, and preceding sub-queries, and is trained using trajectory-relative learning signals derived from the agent's interactions. Across six agent backbones from three model families, ITER consistently outperforms the existing agent-trajectory-trained dense retriever, LRAT, achieving an average relative improvement of 6.9% on InfoSeek-Eval and 15.4% on BrowseComp-Plus. At the matched 4B scale, ITER outperforms AgentIR on InfoSeek-Eval for five of six backbones while achieving a higher visit-to-search recall ratio on BrowseComp-Plus across all six backbones. Ablations further show that structured interaction history and pre-search reasoning provide complementary retrieval context, while previously visited and useful documents, used as redundancy negatives in subsequent searches, provide the strongest trajectory-relative supervision.