通过世界模型扩展自动研究智能体
Scaling Automatic Research Agents via World Models
查看机构详情
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Amazon(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对自动研究智能体扩展时的训练瓶颈,提出WMRL方法,结合两种缓解措施提升收敛性,训练加速3-4倍且性能优于更大规模智能体,还可迁移至多类任务。
中文摘要 AI 辅助
将实证研究自动化是人工智能领域长期以来的研究方向。近期的自动研究(AutoResearch)智能体让这一目标触手可及,因为现代大型语言模型(LLM)具备独立实现解决方案并从执行结果中学习的能力。在这些进展背后,后训练(尤其是强化学习(RL))发挥着核心作用。本文中,我们识别出扩展这类智能体的RL时存在的一个根本矛盾:每条AutoResearch轨迹的两个组成部分(智能体生成与环境执行)的扩展方式截然不同,因为所有生成过程通过批处理共享计算资源,而每次执行则占用其专属的沙箱和实际机器时间。因此,环境执行在训练成本中占主导地位,并随着轨迹数量的增长成为瓶颈。为解决这一矛盾,我们提出世界模型强化学习(WMRL),它用世界模型替代环境执行以消除这一瓶颈。此外,世界模型可能存在缺陷,因为其奖励会受到偏差和噪声的影响。因此,我们进一步为WMRL配备了两种缓解措施:在线去偏和逆方差去噪,分别用于抵消偏差和抑制噪声。从理论上讲,我们证明WMRL的这两种缓解措施严格提升了收敛保证。从实验来看,WMRL在不同智能体规模的各类任务上使训练速度加快了3-4倍,同时超过了标准RL基线的性能。此外,我们经过后训练的40亿参数(4B)和90亿参数(9B)智能体在保留的基准测试上,性能优于大得多的480亿参数(48B)和1200亿参数(120B)开源权重智能体。除AutoResearch外,WMRL还可迁移到后训练的具身视觉语言动作(VLA)策略,这证明了我们方法的通用性。
英文摘要
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes. Behind these gains, post-training (especially RL) plays a central role. In this paper, we identify a fundamental tension when scaling RL for these agents: the two components of every AutoResearch trajectory (agent generation and environment execution) scale in very different manners, since all generation shares compute through batching, while each execution occupies its exclusive sandbox and real machine time. As a result, the environment execution dominates the training cost and becomes the bottleneck as trajectories grow. To resolve this tension, we propose World Model RL (WMRL), which replaces environment execution with a world model to remove this bottleneck. Additionally, the world model can be imperfect, as its rewards are corrupted by bias and noise. Therefore, we further equip WMRL with two mitigations, Online Debiasing and Inverse-Variance Denoising, which offset the bias and suppress the noise respectively. Theoretically, we prove that both mitigations of WMRL strictly improve the convergence guarantee. Empirically, WMRL accelerates training by 3-4x on various tasks at different agent scales, while exceeding the performance of standard RL baselines. Moreover, our post-trained 4B and 9B agents outperform much larger open-weight agents of 48B and 120B on held-out benchmarks. Beyond AutoResearch, WMRL also transfers to post-training embodied VLA policies, which demonstrates the generalizability of our method.