arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SearchMaster:面向搜索智能体的基于接地与调控的自博弈框架

SearchMaster: Grounded and Regulated Self-Play for Search Agents

Wentao Tan, Qiong Cao, Jiaqi Wang, Nan Duan

arXiv 2608.01822首次发表:更新:

AI 中文总结

本文提出SearchMaster自博弈框架,通过ECG、SDR、OOP三项调控措施优化LLM搜索智能体训练,在六个深度搜索基准上显著提升Qwen3.5-9B模型的平均准确率,无需人工标注数据。

AI 中文摘要

训练基于大语言模型(LLM)的搜索智能体需要高质量搜索数据,这类数据需包含要求真实多跳检索的任务,以及能有效使用搜索工具的轨迹。现有流程通常依赖人工编写的任务、专家演示或更强的教师模型。本文提出SearchMaster,这是一种自博弈框架,可在本地搜索环境中,利用自身生成、解决并验证的搜索任务来训练单个LLM。核心挑战在于,自生成的任务和轨迹可能产生误导性信号:伪多跳问题、忽略搜索深度的成功率难度估计,以及过度打开文档但几乎没有针对性证据获取的轨迹。SearchMaster通过三项调控措施解决这些失败模式:一是证据链生成器(ECG),将任务生成基于明确的跨文档证据链,以减少伪多跳问题;二是搜索深度奖励(SDR),通过成功轨迹的搜索深度而非仅成功率来评估任务难度,保留搜索密集型任务;三是过度打开惩罚(OOP),通过抑制过度打开文档来调控工具使用,避免冗长但浅层的浏览。验证后的提议器与求解器轨迹随后通过GRPO联合优化。在六个深度搜索基准测试中,SearchMaster将Qwen3.5-9B骨干模型的平均准确率从38.19%提升至51.52%,在BrowseComp-Plus上的准确率提升达30.1个百分点。这些结果表明,基于接地与调控的自博弈可提供有效的搜索智能体训练数据,无需人工标注的问答对或专家演示。代码可在此https URL获取。

英文摘要

Training LLM-based search agents requires high-quality search data: tasks that demand genuine multi-hop retrieval and trajectories that use search tools effectively. Existing pipelines often depend on human-written tasks, expert demonstrations, or stronger teacher models. We present SearchMaster, a self-play framework that trains a single LLM from search tasks it generates, solves, and verifies in a local search environment. The key challenge is that self-generated tasks and rollouts can yield misleading signals: pseudo multi-hop questions, success-rate difficulty estimates that ignore search depth, and rollouts with excessive opening but little targeted evidence acquisition. SearchMaster addresses these failure modes with three controls. An Evidence-Chain Generator (ECG) grounds task generation in explicit cross-document evidence chains to reduce pseudo multi-hop questions. A Search-Depth Reward (SDR) scores task difficulty by the search depth of successful rollouts rather than success rate alone, keeping retained tasks search-intensive. An Over-Opening Penalty (OOP) regulates tool use by discouraging excessive document opening, avoiding long but shallow browsing. Verified Proposer and Solver rollouts are then jointly optimized with GRPO. Across six deep-search benchmarks, SearchMaster improves a Qwen3.5-9B backbone from 38.19% to 51.52% average accuracy, with a 30.1-point gain on BrowseComp-Plus. These results show that grounded and regulated self-play can provide effective search-agent training data without human-labeled QA pairs or expert demonstrations. The code is available at https://github.com/WentaoTan/SearchMaster.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑