arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作为AI研究世界模型的语言模型

Language Models as AI Research World Models

Zijun Wang, Zewen Liu, Minhua Lin, Zhaotian Weng, Zhan Shi, Bing He, Yisi Sang, Dakuo Wang, Benoit Dumoulin, Wei Jin, Yuyin Zhou, Cihang Xie, Hanqing Lu

arXiv 2610.12235首次发表:更新:

发表机构

Amazon; UC Santa Cruz; Emory University; Pennsylvania State University; UC Santa Barbara(亚马逊; 加州大学圣克鲁兹分校; 埃默里大学; 宾夕法尼亚州立大学; 加州大学圣巴巴拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出将语言模型作为研究世界模型(RWM),利用2600余条实验记录训练的RWM可提升干预结果预测性能,在Qwen3环境中降低78%选择遗憾,还能提升多轮自动研究的最佳收益。

AI 中文摘要

AI研究智能体可自动化提出、实施与评估实验的循环,为递归自我改进开辟了路径。然而其提出实验的能力远超在真实环境中执行的能力,使得结果预测成为有限实验预算下持续自我改进的关键能力。我们将语言模型作为研究世界模型(Research World Models, RWMs),用于预测研究环境中候选干预措施的结果。我们的评估基于来自9个研究环境的2600多条实验记录,涵盖预训练、后训练与推理阶段,对应超过171000个H100 GPU小时的实验量。从真实实验经验中获取的研究知识可提升RWM对同一环境中未见干预措施的预测性能(斯皮尔曼相关系数提升0.27),且可跨环境复用。例如,仅使用来自OLMo3、Marin和Nanochat的预训练经验,RWM在Qwen3环境中相比零经验设置将选择遗憾降低78%。这些益处延伸至固定选择预算下的多轮自动研究:具备环境内与跨环境研究知识的RWMs,分别将最佳收益提升15.8%与11.6%。对13种用作RWMs的语言模型的消融实验显示,添加研究知识相比仅改变模型或增加推理工作量,能更有效地改进干预措施排序。这些发现支持语言模型作为RWMs的可行性,并为未来RWM训练积累实验数据提供了动机。

英文摘要

AI research agents automate the cycle of proposing, implementing, and evaluating experiments, opening a path toward recursive self-improvement. Yet their ability to propose experiments outpaces their capacity to execute them in real environments, making outcome prediction a key capability for sustained self-improvement under limited experimental budgets. We investigate language models as Research World Models (RWMs), which predict the outcomes of candidate interventions across research environments. Our evaluation draws on over 2,600 experimental records from nine research environments spanning pretraining, post-training, and inference, representing more than 171,000 H100 GPU-hours of experimentation. Research knowledge acquired from real experimental experience improves RWM predictions of unseen interventions within the same environment (Spearman +0.27), and can be reused across environments. For example, using only pretraining experience from OLMo3, Marin, and Nanochat, an RWM reduces selection regret in the Qwen3 environment by 78% compared with zero-experience setting. These benefits extend to multi-round Autoresearch under a fixed selection budget: RWMs with in-env and cross-env research knowledge increase the best gain achieved by 15.8% and 11.6%, respectively. Ablations across 13 language models used as RWMs show that adding research knowledge can improve intervention ranking more than changing models or increasing reasoning effort alone. These findings support language models as RWMs and motivate accumulating experimental data for future RWM training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑