arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于预测导航的深度研究预训练

Deep Research Pretraining via Predictive Navigation

Jiang Zhou, Zhiyuan Fan, Xing Wu, Tinghao Yu, Feng Zhang, Lilin Wang

arXiv 2608.00432首次发表:更新:

发表机构

Tencent; Hunyuan Team(腾讯; 混元团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Deep Research Pretraining(DRP)离线框架,在学术引用图和维基百科超链接上预训练Qwen3-14B-Base模型,可提升智能体在多个基准任务上的表现,是轨迹智能体训练的补充方法。

AI 中文摘要

深度研究智能体通常在昂贵的、与环境绑定的工具使用轨迹上进行训练,这些轨迹需要重复检索、文档检查和报告评估。我们提出Deep Research Pretraining(DRP),这是一种离线框架,可从自然存在的证据结构中推导预测导航监督。给定带有引用或超链接的段落,DRP构建代理研究目标,恢复链接证据和图相关替代方案,并将其转换为搜索-打开-撰写轨迹。这教会模型要搜索什么、检查哪些文档以及如何综合证据,无需实时检索环境或执行策略滚动。我们在学术引用图(DRP-Paper)和维基百科超链接(DRP-Web)上实例化DRP,在1B token上持续预训练独立的Qwen3-14B-Base模型,并在13K智能体轨迹的受控部分上微调它们。在每个低数据预算下的五个独立采样子集上,两种变体在DeepResearch Bench上始终优于匹配的无DRP模型。使用四分之一的SFT数据,DRP-Web甚至超过了固定的无DRP全数据检查点,收益可迁移到ResearchQA、WebWalkerQA和SimpleQA。从匹配的低数据SFT检查点开始,DRP-Web的优势在后续的智能体强化学习中也持续存在。源匹配和证据不匹配控制表明,这些改进源于证据条件导航,而非领域暴露或智能体格式模仿。因此,DRP为基于轨迹的智能体训练提供了一种有前景的补充方法。

英文摘要

Deep research agents are often trained on expensive, environment-grounded tool-use trajectories that require repeated retrieval, document inspection, and report evaluation. We introduce Deep Research Pretraining (DRP), an offline framework that derives predictive navigation supervision from naturally occurring evidence structures. Given a citation-bearing or hyperlinked passage, DRP constructs a proxy research objective, recovers linked evidence and graph-related alternatives, and converts them into search-open-write trajectories. This teaches models what to search for, which documents to inspect, and how to synthesize evidence, without a live retrieval environment or executed policy rollout. We instantiate DRP on scholarly citation graphs (DRP-Paper) and Wikipedia hyperlinks (DRP-Web), continually pretrain separate Qwen3-14B-Base models on 1B tokens, and fine-tune them on controlled fractions of 13K agent trajectories. Across five independently sampled subsets at each low-data budget, both variants consistently outperform matched no-DRP models on DeepResearch Bench. With one quarter of the SFT data, DRP-Web even surpasses a fixed no-DRP full-data checkpoint, with gains transferring to ResearchQA, WebWalkerQA, and SimpleQA. Starting from matched low-data SFT checkpoints, the DRP-Web advantage also persists through subsequent agentic RL. Source-matched and evidence-mismatch controls indicate that these improvements arise from evidence-conditioned navigation rather than domain exposure or agent-format imitation. DRP thus provides a promising complementary approach to trajectory-based agent training.

Commentsworking in progress; correspondence to {ucaswu,maxwellyu}@tencent.com or wuxing@iie.ac.cn

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑