arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AutoWorldModel-Bench:面向自动化世界模型研究的以状态为中心的基准

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri

arXiv 2608.11216首次发表:更新:

发表机构

Electronic Arts; Simon Fraser University(美国艺电公司; 西蒙菲莎大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AutoWorldModel-Bench是面向AI编码智能体的闭环基准,涵盖8个游戏环境,采用结构化状态表征,64次会话中多数智能体通过研究式修改改进了世界模型,可评估智能体的开放式研究能力。

AI 中文摘要

世界建模是一个尚未解决的领域:架构、训练目标和状态表征以复杂方式相互作用,且没有单一方案能在所有环境中占据主导地位。这使其成为作为自主研究者的AI编码智能体的理想测试平台——在此场景中,改进方向未预先指定,与当前智能体基准中占主导的“按规范工程实现”任务不同。我们推出AutoWorldModel-Bench,这是一个闭环基准,前沿编码智能体在固定计算预算下自主改进提供的世界模型 starter。该基准涵盖8个游戏环境,采用统一结构化状态表征——从每个游戏中提取的真实实体状态,通过共享张量格式使用——这将动力学建模与感知分离,使每次运行迭代仅需数分钟。在64次会话中,Codex-5.4和Claude Opus 4.6在63次会话中改进了其starter;在91%的会话中,获胜的编辑是一项非平凡的研究式修改——新目标、表征、回滚程序或架构变更——而非超参数调整。我们的基准提供了一个场景,可在其中对前沿编码智能体进行开放式研究而非按规范工程实现问题的评估。

英文摘要

World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setting in which the improvement direction is not specified in advance, unlike the engineering-to-spec tasks that dominate current agent benchmarks. We introduce AutoWorldModel-Bench, a closed-loop benchmark in which frontier coding agents autonomously improve a provided base world model under a fixed compute budget. The benchmark spans eight game environments under a unified structured-state representation--ground-truth entity state extracted from each game and consumed through a shared tensor format--which isolates dynamics modeling from perception and enables minutes-per-run iteration. Across 64 sessions, Codex-5.4 and Claude Opus 4.6 improve their base on a held-out test split in all but one session, with about half (33 of 64) a substantial gain ($Δ\geq +0.10$) and the remaining improvements smaller but positive; in 91% of sessions the winning edit is a substantive change to the model or training rather than a hyperparameter tweak. Our benchmark offers a setting in which frontier coding agents can be evaluated on open-ended research rather than engineering-to-spec problems.

CommentsProject page: https://electronicarts.github.io/AutoWorldModelBench/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑