发表机构
University College London; University of Liverpool(伦敦大学学院; 利物浦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出Memento 3,让冻结LLM智能体通过外部记忆学习显式世界模型,经循环优化后在ARC-AGI-3和Atari Pong任务中取得优异表现,实现模型化递归自我改进。
AI 中文摘要
在陌生环境中学习行动需要智能体推断世界的运行规律,并在新证据出现时修正这一认知。然而有限的观测结果可能支撑多种世界模型,这些模型既能解释过往交互,却会在未见过的状态下预测出不同结果。我们推出Memento 3,它基于Memento系列,让冻结的大语言模型(LLM)智能体能通过外部记忆持续学习显式世界模型。该智能体将自然语言规则手册作为持久语义记忆,记录可修正的环境动态假设,同时对未知部分不作明确规定;它会把该规则手册编译为可执行代码,用于预测与规划。通过观测、反思、规则修正、编译和验证的持续循环,智能体利用预测误差优化规则手册及其代码,仅当LLM判定更新后的代码符合规则手册,且细胞级精确重放能复现观测到的状态转移时,才会接受该代码。我们将这一过程视为递归自我改进(RSI)的模型化路径:智能体自主探索环境、修正世界模型,并利用经验证的更新指导后续交互与学习,而底层LLM保持固定。种群扩展机制会并行维护多个世界模型,共享交互证据并利用其预测指导探索。在ARC-AGI-3上,单模型智能体通过了全部25款公开游戏的所有关卡,达到平均人类行动相对效率(RHAE)100.0,仅使用人类行动数的44%。在Atari Pong案例研究中,学习到的反馈控制器在3次不同开局的评估回合中均以21:0获胜,且无需额外LLM调用。
英文摘要
Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.