发表机构
HKUST; Tencent; HKUST(GZ)(香港科技大学; 腾讯; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究训练工具使用智能体从自身经验改进的难题,提出基于深度搜索世界的自蒸馏框架深度搜索进化,通过迭代操作训练智能体,该框架无需更强大模型蒸馏,性能有竞争力,还将发布相关资源促进后续研究。
AI 中文摘要
训练工具使用智能体从自身经验中改进仍具有挑战性,监督微调依赖固定教师蒸馏轨迹,稀疏奖励强化学习对长期交互的监督较弱。我们提出了深度搜索进化(DeepSearch-Evolve),这是一个基于深度搜索世界(DeepSearch-World)构建的网络智能体自蒸馏框架,它是一个具有可重现搜索和页面读取工具的确定性和可验证环境。深度搜索世界包含由实体级随机游走构建的42万个多跳问答任务,并支持对自我进化有用的关键智能认知行为。深度搜索进化迭代执行轨迹生成、过滤、数据混合和微调以训练更强的智能体。在没有从更强大模型进行蒸馏情况下,深度搜索世界-9B与开源智能体相比具有竞争力,在BrowseComp上达到31.2%,在GAIA上达到61.5%,在HotPotQA上达到93.4%,表明可验证环境能实现长期网络智能体的可扩展自我进化。我们将发布环境、42万个训练池、验证集、模型和代码以促进未来对自我改进深度搜索智能体的研究。
英文摘要
Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak supervision for long-horizon interactions. We present DeepSearch-Evolve, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and verifiable environment with reproducible search and page-reading tools. DeepSearch-World contains 420K multi-hop QA tasks constructed from entity-level random walks and supports key agentic cognitive behaviors useful for self-evolving, including progress verification, grounded reflection, and failure recovery. DeepSearch-Evolve iteratively performs trajectory generation, filtering, data mixing, and fine-tuning to train stronger agents. Without distillation from more capable models, DeepSearch-World-9B achieves competitive performance compared with open-source agents, reaching 31.2% on BrowseComp, 61.5% on GAIA, and 93.4% on HotpotQA, showing that verifiable environments enable scalable self-evolution for long-horizon web agents. We will release the environment, 420K training pool, validation set, model, and code to facilitate future research on self-improving deep search agents.