AI 中文总结
针对深度研究智能体长上下文理解不足的问题,提出DR-to-Long方法将DR-RL轨迹转为长上下文QA数据,并构建DLD-RL训练流程,在深度研究和长上下文基准上分别提升7.3%和13.5%。
AI 中文摘要
深度研究(DR)智能体通过多轮搜索和访问与现实世界的网络环境交互,导致其上下文随时间快速增长。我们观察到,即使在DR智能体强化学习(DR-RL)之后,模型剩余的预测错误中仍有61.6%可归因于长上下文理解不足,包括长上下文幻觉和跨文档证据整合失败。这促使我们通过增强模型的长上下文能力来进一步突破DR-RL的瓶颈。然而,有效的长上下文训练不仅仅需要增加上下文长度。为弥补数据缺口,我们提出了“DR滚动轨迹到长上下文问答(DR-to-Long)”方法。该方法重新利用DR-RL轨迹,这些轨迹天然包含搜索历史、访问过的网页、证据片段和最终答案监督。然后,它将每个轨迹中的紧凑片段和网页摘要替换为对应URL的完整内容,从而生成更长的多文档上下文,同时保留原始证据关系。基于DR-to-Long,我们引入了DLD(DR -> LongQA -> DR)-RL。DLD-RL首先执行一个短暂的DR-RL阶段以收集滚动轨迹,然后以零注释成本将这些轨迹转换为LongQA实例。随后,模型通过LongQA-RL进行优化以增强长上下文能力,接着进行完整的DR-RL以继续提升其DR能力。实验表明,DLD-RL在三个深度研究基准上比标准DR-RL高出7.3%,并在三个长上下文基准上提升了13.5%的性能。
英文摘要
Deepresearch (DR) agents interact with real-world web environments through multi-turn search and visit, causing their contexts to grow rapidly over time. We observe that, even after DR Agentic Reinforcement Learning (DR-RL), 61.6% of the model's remaining prediction errors can still be attributed to insufficient long-context understanding, including longcontext hallucination and failures in cross-document evidence integration. It motivates us to further break the bottleneck of DR-RL by strengthening the model's long-context ability. However, effective LongContext training requires more than simply increasing context length. To bridge the data gap, we propose `DR Rollouts to LongContext-QA (DR-to-Long)'. The method repurposes DR-RL trajectories, which naturally contain search histories, visited webpages, evidence snippets, and final-answer supervision. It then replaces the compact snippets and webpage summaries in each trajectory with the full contents of their corresponding URLs, producing substantially longer multi-document contexts while preserving the original evidence relationships. Building on DR-to-Long, we introduce DLD (DR -> LongQA -> DR)-RL. DLD-RL first performs a short DR-RL stage to collect rollout trajectories, which are then converted into LongQA instances at zero annotation cost. The model is subsequently optimized with LongQA-RL to strengthen LongContext ability, followed by full DR-RL to continue improving its DR capability. Experiments show that DLD-RL outperforms standard DR-RL by 7.3% on three Deepresearch benchmarks and improves performance by 13.5% on three long-context benchmarks.