发表机构
Xiaohongshu Inc.; the Joint SDU-NTU Centre for Artificial Intelligence Research (C-FAIR), Shandong University; University of Illinois at Urbana-Champaign(小红书公司; 山东大学-南洋理工大学联合人工智能研究中心(C-FAIR); 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Harness-Search通过多智能体协调解耦检索、状态更新和终止决策,形成提议-提交-审计循环,在七个长时程搜索基准上显著提升检索与答案生成性能。
AI 中文摘要
长时程搜索要求智能体跨多个步骤收集证据,并将其综合为有充分依据的答案。最近的智能体框架(agent harnesses)为支持此类长期运行的搜索过程提供了一个自然且有前景的框架。随着交互历史的增长,框架中的单个智能体可能会陷入困境,导致策略丢失未解决问题的线索、忽略有用的证据,或在收集到足够支持之前终止。一种有前景的方法是将提出检索行动、更新持久状态以及决定何时停止这三项不同职责解耦,而不是将它们集中在一个单一策略中。针对这一点,我们引入了Harness-Search,一种多智能体搜索框架,以减少在后续探索、证据整理和终止决策中传播的局部错误。具体而言,Harness-Search将这些职责分配给三个权限受限的权威机构:一个检索策略(Retrieval Policy),负责提出搜索操作;一个记忆操作器(Memory Operator),负责验证并提交持久状态更新;以及一个摘要审计器(Summary Auditor),根据整理证据的充分性接受或拒绝终止。这些角色共同形成了一个提议-提交-审计(Propose-Commit-Audit)循环,在该循环中,行动被提议,持久证据被选择性提交,停止决策受到明确的充分性检查。在七个长时程搜索基准测试中,Harness-Search在相同的策略骨干下同时提高了检索和答案生成性能,在每个证据检索基准上,与最强的基于框架的基线相比,召回率提高了4.60-27.92个百分点,最终答案召回率提高了12.34-30.13个百分点。此外,轨迹级分析表明,随着搜索历史的增长,Harness-Search持续积累有用证据并扩大证据覆盖范围,同时减少冗余检索。
英文摘要
Long-horizon search requires agents to gather evidence across multiple steps and synthesize it into well-supported answers. The recent agent harnesses provide a natural and promising framework to support such long-running search processes. As interaction histories grow, one single agent in harnesses might get stuck and cause the policy to lose track of unresolved questions, overlook useful evidence, or terminate before sufficient support has been collected. One of promising way is to decouple three distinct responsibilities of proposing retrieval actions, updating persistent state, and deciding when to stop rather than concentrating them within a single policy. Targeted at it, we introduce Harness-Search, a multi-agent search harness to reduce the local errors propagating across subsequent exploration, evidence curation, and termination decisions. In particular, Harness-Search assigns these responsibilities to three permission-bounded authorities: a Retrieval Policy that proposes search operations, a Memory Operator that validates and commits persistent-state updates, and a Summary Auditor that accepts or rejects termination based on the sufficiency of the curated evidence. Together, these roles form a Propose-Commit-Audit loop in which actions are proposed, persistent evidence is selectively committed, and stopping decisions are subjected to an explicit sufficiency check. Across seven long-horizon search benchmarks, Harness-Search improves both retrieval and answer generation under the same policy backbone, increasing Recall by 4.60-27.92 points and Final-Answer Recall by 12.34-30.13 points over the strongest harness-based baseline on each evidence-retrieval benchmark. Moreover, trajectory-level analyses show that Harness-Search continues to accumulate useful evidence and expand evidence coverage with less redundant retrieval as the search history grows.
Comments26 pages, Natural Language Processing