SearchArt:使用可扩展的合成和经过验证的任务训练长视野搜索代理
SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task
- Huawei Cloud(华为云)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对训练长视野搜索代理的挑战,SearchArt提出通过验证驱动任务合成及多阶段训练管道的框架,构建大规模数据集并验证数据可靠性,经此训练的代理在多基准测试中表现出色,参数少却成绩优异。
AI中文摘要:
大语言模型的进展使搜索代理能自主处理复杂任务,但训练有效的搜索代理仍具挑战。我们引入SearchArt,这是一个通过验证驱动的任务合成和多阶段训练管道来训练长视野搜索代理的可扩展框架。它通过合成问答对和搜索轨迹构建大规模数据集,并设计验证管道确保数据可靠性。经验证的轨迹用于多阶段训练。实验表明,仅27B参数的SearchArt在多个基准测试中得分优异,匹配或超越前沿闭源代理。
英文摘要:
Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning horizons. However, training effective search agents remains challenging due to the lack of scalable and long-horizon tasks, and the difficulty of evaluating and correcting intermediate reasoning and tool-use behaviors. We introduce SearchArt, a scalable framework for training long-horizon search agents through verification-driven task synthesis and a multi-stage post-training pipeline. SearchArt constructs large-scale datasets for complex search-, research- and user-oriented tasks by synthesizing diverse information-seeking QA pairs and corresponding search trajectories from web documents and automatically generated evidence graphs. To ensure the reliability of the synthesized data, we design a verification pipeline that jointly evaluates QA consistency, trajectory quality, and the relevance of retrieved evidence. The verified trajectories are subsequently used in a multi-stage training process comprising supervised fine-tuning and reinforcement learning-based policy optimization. Search agents trained with SearchArt exhibit adaptive search planning, iterative evidence aggregation, and complex reasoning over extended interaction horizons. Experimental results demonstrate that, with only (Qwen3.5-) 27B parameters, SearchArt scores 74.39 on BrowseComp-ZH, 70.06 on BrowseComp, and 52.55 on Deepresearch-bench, matching or surpassing frontier closed-source agents on both deepsearch and deepresearch benchmarks.