AI 中文总结
针对深度搜索智能体在长视界下因检索敏感而失败的问题,提出LexiHorizon框架,通过扩展上下文预算、窗口化管理检索内容并保留推理历史,以及引入结果门控搜索努力奖励,在9B模型上显著优于基线,最大绝对增益达23.8个百分点。
AI 中文摘要
深度搜索智能体通过迭代检索、多跳推理和跨多源证据综合来处理复杂知识任务。现有方法通常假设检索系统相对稳定,并在短视界工具交互中运行。然而,当检索对查询表述敏感时,即使语义上恰当的查询也可能因实体名称、别名或关键词组合不匹配而无法呈现关键证据。从此类失败中恢复需要重复的查询重述和更长的交互轨迹。这种设置带来了独特的训练挑战,因为策略必须在管理不断扩大的检索内容量的同时,维持长视界的查询探索。我们提出LexiHorizon,一个用于在长视界上训练搜索智能体的框架,它扩展了轨迹上下文预算,使用近期工具观察的窗口来管理累积的检索内容,同时保留推理历史,并引入一个以结果门控的搜索努力奖励,该奖励为具有非零答案奖励的轨迹的工具调用提供有界奖励。在XBench、WebWalkerQA和BrowseComp-ZH上的实验表明,由此产生的9B模型持续优于其基础模型和MiroThinker-1.7-mini,最大绝对增益分别为8.7和23.8个百分点。这些结果表明,将扩展的上下文预算与保留推理的上下文管理相结合,有利于长视界深度搜索智能体。
英文摘要
Deep search agents tackle complex knowledge tasks through iterative retrieval, multi-hop reasoning, and evidence synthesis across multiple sources. Existing approaches typically assume relatively stable retrieval systems and operate over short-horizon tool interaction. However, when retrieval is sensitive to query formulation, even a semantically appropriate query may fail to surface critical evidence because of mismatched entity names, aliases, or keyword combinations. Recovering from such failures requires repeated query reformulation and longer interaction trajectories. This setting poses a distinct training challenge, as the policy must sustain long-horizon query exploration while managing an expanding volume of retrieved content. We propose LexiHorizon, a framework for training search agents over long horizons that expands the trajectory context budget, manages accumulated retrieval content using a window over recent tool observations while preserving the reasoning history, and introduces an outcome-gated search-effort reward that provides a bounded bonus for tool invocations to trajectories with nonzero answer reward. Experiments on XBench, WebWalkerQA, and BrowseComp-ZH show that the resulting 9B model consistently outperforms both its base model and MiroThinker-1.7-mini, with maximum absolute gains of 8.7 and 23.8 percentage points, respectively. These results suggest that combining an extended context budget with reasoning-preserving context management benefits long-horizon deep search agents.