AI 中文总结
本研究通过替换预测目标与随机奖励,审计世界模型后训练,发现性能提升部分源于优化过程而非纯预测,且随机奖励仍能提升任务覆盖率。
AI 中文摘要
在行动之前预测环境将如何变化,是智能体做出更好决策的自然途径。因此,近期的后训练方法要求智能体预测下一个观测,并将该预测转化为奖励或直接监督信号,这被称为世界模型。现有的下一观测训练方法帮助智能体学习环境内容。然而,它们还涉及一个优化过程,这可能引入除学习预测世界之外的多重效应。因此,训练过程中性能提升的来源仍是一个开放问题。我们通过用分布内不匹配的观测替换真实下一观测目标来回答这一研究问题。在两个交互式文本环境中,与真实目标相比,不匹配目标将预测准确率降低了15.3%-61.6%,但仍保留了相对于基础模型的显著任务收益。与基础模型相比,训练后的模型考虑更多候选动作,并表现出更少的循环。我们还引入了一种设置,用独立的随机信号替换基于预测的奖励。这种训练扩展了任务覆盖率(pass@64),即使奖励不携带环境信息。我们还将这一发现推广到VisualWebArena,在随机奖励训练下,相对于基础模型,pass@64提高了14.3%,而无需观测匹配奖励或外部多模态教师来构建奖励。
英文摘要
Predicting how an environment will change before acting is a natural route to better decision making for agents. Recent post-training methods therefore require agents to predict the next observation and turn that prediction into a reward or a direct supervision signal, which is called world model. Existing next-observation training methods help the agent to learn the environmental content. However, they additionally involve an optimization process, which may introduce several effects other than learning to predict the world. Consequently, where the performance gain comes from during the training process remains an open question. We answer this research question through replacing true next-observation targets with in-distribution mismatched observations during the training process. Across two interactive text environments, mismatched targets lower prediction accuracy by 15.3-61.6% relative to ground-truth targets, yet retain substantial task gains over the base model. Compared with the base model, trained models consider more candidate actions and exhibit less looping. We also introduce a setting that replaces prediction-based rewards with independent random signals. This training expands task coverage (pass@64) even when the reward carries no environment information. We also generalize this finding to VisualWebArena, where random-reward training raises pass@64 by 14.3% relative to the base model, without observation-matching rewards or an external multimodal teacher for reward construction.
Comments24 pages, 6 figures. Preprint