IterSynth:通过角色解耦的迭代综合重新思考深度搜索智能体
IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis
浏览论文内容
中文总结 AI 辅助
IterSynth通过角色解耦的迭代综合范式,分离规划与综合,利用摘要作为持久状态,并采用RDPO强化学习,在深度搜索基准上超越现有≤8B智能体,并提升前沿模型的零样本性能。
中文摘要 AI 辅助
深度搜索要求大语言模型智能体分解复杂查询、搜索证据并综合有依据的答案,然而现有的ReAct风格智能体存在两个局限性:角色耦合,即单一策略必须同时处理规划、证据使用和综合;以及上下文累积,即不断增长的搜索历史引入噪声并掩盖有用信息。为解决这些问题,我们提出IterSynth,一种角色解耦且基于摘要的范式,在识别信息需求的规划器与将证据整合到不断演进的摘要状态中的综合器之间交替进行。这种设计将规划与综合分离,同时使用摘要作为搜索的持久状态,从而减少能力耦合和上下文噪声。为有效训练IterSynth,我们进一步引入用于强化学习的角色解耦策略优化(RDPO),它将终端结果奖励与回合级评分标准相结合,并计算角色特定的优势以实现更精确的信用分配。在五个长时程深度搜索基准(如BrowseComp和Xbench-DS)上的实验表明,IterSynth-8B的平均得分为50.7,超过最强的前代≤8B智能体+4.2%。此外,IterSynth作为一种模型无关的提示范式,在前沿专有模型上相比ReAct及类似提示范式带来了显著的零样本提升。
英文摘要
Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories introduce noise and obscure useful information. To address these issues, we propose IterSynth, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information needs and a Synthesizer for integrating evidence into an evolving summary state. This design separates planning from synthesis while using the summary as the persistent state of search, reducing both capability coupling and context noise. To train IterSynth effectively, we further introduce Role-Decoupled Policy Optimization (RDPO) for reinforcement learning, which combines terminal outcome rewards with turn-level rubric evaluations and computes role-specific advantages for more precise credit assignment. Experiments on five long-horizon deep-search benchmarks such as BrowseComp and Xbench-DS show that IterSynth-8B achieves an average score of 50.7, surpassing the strongest prior $\leq$8B agent by +4.2\%. Moreover, IterSynth serves as a model-agnostic prompting paradigm, delivering substantial zero-shot gains over ReAct and similar prompting paradigms on frontier proprietary models.
发表机构
- Zhejiang University(浙江大学)
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。