AI 中文总结
针对多模态深度研究智能体在长上下文(128k)和长交互(75+回合)下的强化学习训练难题,提出Long-MDR三组件训练方案,显著提升学习效率与稳定性,并在50回合评估中于多数基准上取得领先。
AI 中文摘要
下一代多模态研究智能体必须对长期的研究历史进行推理,而非短期的模型补全。在单个任务中,智能体可能反复搜索网络、检查视觉证据、重新审视早期假设,并积累数万token的多模态上下文。尽管存在这一趋势,针对多模态研究智能体的在线强化学习仍主要局限于较短的上下文和交互范围。我们将在线强化学习训练扩展到128k上下文和75次以上的工具交互回合。据我们所知,这是首次在128k上下文下进行的在线多模态深度研究强化学习研究,也是首次在75个工具回合范围内进行训练的研究。扩展到这一范围暴露了传统强化学习训练的若干实际限制。在训练早期,弱策略无法有效利用大型交互预算,导致昂贵的展开但奖励提升甚微。随后,策略熵可能在性能饱和之前崩溃,过早终止有用的学习。我们提出了Long-MDR,一种专门针对该场景设计的三组件训练方案:在线策略蒸馏预热、渐进式范围扩展和熵触发救援。这些技术共同提高了长范围强化学习的学习效率和稳定性,使得在直接训练缓慢且昂贵的场景中能够持续取得进展。在50回合评估预算下,我们经过强化学习训练的Long-MDR-9B在六个基准中的五个上,在7B-9B智能体比较中排名第一。
英文摘要
The next generation of multimodal research agents must reason over long-lived research histories rather than short model completions. During a single task, an agent may repeatedly search the web, inspect visual evidence, revisit earlier hypotheses, and accumulate tens of thousands of tokens of multimodal context. Despite this trend, online RL for multimodal research agents remains largely confined to shorter contexts and interaction horizons. We push online RL training to 128k context and 75+ tool-interaction turns. To our knowledge, this is the first online multimodal deep-research RL study trained at 128k context, and the first trained with a 75 tool-turn horizon. Scaling to this regime exposes several practical limitations of conventional RL training. Early in training, weak policies make poor use of large interaction budgets, causing expensive rollouts with little reward improvement. Later, policy entropy can collapse before performance has saturated, prematurely ending useful learning. We introduce Long-MDR, a three-component training recipe designed specifically for this setting: On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rescue. Together, these techniques improve both the learning efficiency and stability of long-horizon RL, enabling continued gains in a regime where direct training is slow and costly. At a 50-turn evaluation budget, our RL-trained Long-MDR-9B ranks first on five of six benchmarks among the compared 7B-9B agents.