TRACE:搜索智能体并行扩展的轨迹选择
TRACE: Trajectory Selection for Parallel Scaling of Search Agents
浏览论文内容
中文总结 AI 辅助
TRACE提出基于跨轨迹证据的轻量级选择器,通过排序完整轨迹替代投票或生成式聚合,提升并行搜索智能体的答案选择准确率与效率。
中文摘要 AI 辅助
并行搜索可能生成正确答案,但最终答案投票可能无法选出该答案。我们将这一整合阶段形式化为轨迹选择,并引入TRACE(基于聚合跨轨迹证据的轨迹排序),这是一个轻量级的学习型选择器,利用答案背后的搜索证据对完整轨迹进行排序。TRACE保留单个查询和证据的出现信息,通过共享内容或文档身份连接各轨迹,并跨这些关系传播信息。每个候选答案随后读取其自身轨迹的更新状态,在保留检索来源的同时整合来自相关轨迹的证据。TRACE在冻结文本嵌入上使用答案级监督进行训练,无需额外搜索或自回归聚合即可返回现有答案。每个搜索设置一个选择器即可跨轨迹策略和智能体骨干迁移,无需针对特定智能体进行微调,在K=16时,在六个WebQA策略和六个长时程数据集-骨干组合上优于投票方法。在Qwen2.5-14B Base/SFT WebQA池上,TRACE达到45.2/49.2%的EM分数,而最强的Qwen3-32B生成式聚合器为43.9/48.0%。在长时程FRAMES、GAIA和BrowseComp上,TRACE达到78.6%的平均准确率,超过多数投票3.1个百分点。在Base WebQA池上,仅使用8条轨迹的TRACE与使用64条轨迹的多数投票相比,差距在0.4个百分点以内。TRACE在所有七个WebQA基准上的处理吞吐量至少比SolAgg、SummAgg和AggAgent高10倍。这些结果表明,重用跨轨迹搜索证据为并行搜索提供了一种有效且高效的替代重型生成式聚合的方法。代码可在以下URL获取。
英文摘要
Parallel search may generate a correct answer that final-answer voting fails to select. We formulate this consolidation stage as trajectory selection and introduce TRACE (Trajectory Ranking with Aggregated Cross-Rollout Evidence), a lightweight learned selector that ranks completed trajectories using the search evidence behind their answers. TRACE preserves individual query and evidence occurrences, connects rollouts through shared content or document identity, and propagates information across these relations. Each candidate answer then reads the updated states of its own trajectory, preserving retrieval provenance while incorporating evidence from related rollouts. Trained with answer-level supervision over frozen text embeddings, TRACE returns an existing answer without additional search or autoregressive aggregation. One selector per search setting transfers across rollout policies and agent backbones without agent-specific fine-tuning, improving over voting across six WebQA policies and six long-horizon dataset-backbone combinations at $K=16$. On Qwen2.5-14B Base/SFT WebQA pools, TRACE achieves 45.2/49.2% EM, compared with 43.9/48.0% for the strongest Qwen3-32B generative aggregators. On long-horizon FRAMES, GAIA, and BrowseComp, it reaches 78.6% average accuracy, exceeding majority voting by 3.1 percentage points. On Base WebQA pools, TRACE with only 8 rollouts comes within 0.4 points of majority voting over 64. TRACE also achieves at least $10\times$ higher processing throughput than SolAgg, SummAgg, and AggAgent across all seven WebQA benchmarks. These results show that reusing cross-rollout search evidence provides an effective and efficient alternative to heavyweight generative aggregation for parallel search. Code is available at https://github.com/Jaasssoooonnnnn/TRACE.
发表机构
- New York University Shanghai(上海纽约大学)
- New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。