arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeepResearch 智能体系统

DeepResearch Agent System

Yong Huang, Yulu Huang, for the team Collaboration

arXiv 2607.27562首次发表:更新:

发表机构

Software Copyright Research and Development Document(软件版权研究与开发文档)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DeepResearch 智能体系统是基于稀疏激活架构的大语言模型,通过双模式推理引擎、多工具协调等技术,在多个智能体基准测试中取得最优性能,已完全开源并支持多领域应用。

AI 中文摘要

DeepResearch 智能体系统是专为深度信息检索、多步推理和自主研究任务设计的大语言模型系统。该系统基于稀疏激活架构构建,总参数达300亿,每个 token 仅激活30亿参数,在多个智能体搜索基准上实现了最优性能,同时与同等规模的密集模型相比,推理速度提升3.2倍。系统支持128K token 的上下文窗口,其分层注意力机制较标准长上下文方法实现了18.7%的准确率提升和23.4%的召回率提升。双模式推理引擎既提供用于基础多步问题解决的 ReAct 范式,也提供用于高性能迭代研究的 IterResearch 模式,最多支持20个推理步骤,整体较单次通过基线实现了31.2%的准确率提升。多工具协调集成了检索、计算、网页搜索和文件解析模块,工具使用准确率达92.1%。基于 GRPO 算法的强化学习优化框架提供 token 级策略梯度,使训练稳定性提升35%,收敛速度加快42%。带有种子扩展的自动数据合成流水线实现了92.5%的可用性率。基准测试结果包括在 Humanity's Last Exam 上得分87.3%、在 BrowserComp Chinese 上得分85.3%、在 WebWalkerQA 上得分91.2%。该系统已完全开源,涵盖数据合成、训练和推理代码,支持学术研究、商业分析、研发支持和教育领域的应用。

英文摘要

The DeepResearch Agent System is a large language model system engineered for deep information retrieval, multi-step reasoning, and autonomous research tasks. Built upon a sparse activation architecture with 30 billion total parameters of which only 3 billion are activated per token, the system achieves state-of-the-art performance on multiple agent search benchmarks while delivering 3.2 times faster inference compared to dense counterparts of equivalent scale. The system supports a 128K-token context window with hierarchical attention mechanisms that yield 18.7% accuracy and 23.4% recall improvements over standard long-context approaches. A dual-mode reasoning engine provides both a ReAct paradigm for basic multi-step problem solving and an IterResearch mode for high-performance iterative research with up to 20 reasoning steps, collectively delivering a 31.2% accuracy improvement over single-pass baselines. Multi-tool coordination integrates retrieval, computation, web search, and file parsing modules to achieve 92.1% tool-use accuracy. A reinforcement learning optimization framework based on the GRPO algorithm provides token-level policy gradients that improve training stability by 35% and accelerate convergence by 42%. An automated data synthesis pipeline with seed-based expansion achieves a 92.5% usability rate. Benchmark results include 87.3% on Humanity's Last Exam, 85.3% on BrowserComp Chinese, and 91.2% on WebWalkerQA. The system is fully open-sourced, including data synthesis, training, and inference code, and supports applications in academic research, business analysis, R&D support, and education.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑