arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentGarten:用于智能体演化的代码世界

AgentGarten: Code Worlds for Evolving Agents

Jiawei Chi, Shangchen Miao, Zhiyuan Shi, Kailu Wu, Hanyang Wang, Weiliang Chen, Qiyu Dai, Jinshan Ren, Jun Gao, Mingsheng Long, Yueqi Duan, Jiangran Lyu, Jialong Wu, Fangfu Liu

arXiv 2610.12374首次发表:更新:

发表机构

MirroS Technical Report(MirroS技术报告)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AgentGarten是耦合模拟器与游戏引擎并搭配共享神经渲染器的框架,可构建实时交互式环境,使智能体学习效率大幅提升,仅需4轮学习即可完成任务,助力智能体通过交互经验持续演化。

AI 中文摘要

交互式虚拟世界允许智能体通过探索与交互进行学习。智能体所能学习的内容受限于其训练环境,这类环境必须具备一致性,拥有稳定的状态、规则与动态,同时还需具备真实性,其观测结果需符合真实世界的视觉分布。在多样化的世界中同时满足上述两点仍是一个瓶颈问题。我们提出了AgentGarten,这是一个将模拟器与游戏引擎耦合,并搭配共享神经渲染器以构建实时交互式环境的框架。其模拟后端维护持久的世界状态,并执行程序定义的交互规则,而渲染器则通过通用接口导出的结构化条件生成视觉观测结果。为构建该神经渲染器,我们将预训练视频模型适配至几何条件,通过我们提出的对抗强制(Adversarial Forcing)对其进行蒸馏,并优化推理过程以实现实时交互。对抗强制通过精确重放使历史预填具备可微性,从而让后续预测的损失能够更新渲染器对先前观测结果的编码方式,同时加入真实数据的对抗监督以提升其视觉质量。在AgentGarten中,智能体通过视觉观测感知世界,实时与世界交互,并通过将每一轮经验蒸馏成剧本,供后续智能体继承与优化来实现性能提升。我们的实证研究表明,该框架的学习效率有显著提升,智能体仅需4轮学习即可完成任务,而传统强化学习对应模型则需要数百万轮。由于新的世界可通过代码编写并通过同一接口进行渲染,环境的数量与难度可随智能体一同扩展,这是实现智能体通过交互经验持续演化的重要一步。

英文摘要

Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving both across diverse worlds remains a bottleneck. We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments. Its simulation backends maintain persistent world state and execute program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a common interface. To build the neural renderer, we adapt a pretrained video model to geometry conditions, distill it with our proposed Adversarial Forcing, and optimize inference for real-time interaction. Adversarial Forcing makes history prefilling differentiable through exact replay, so that losses on later predictions update how the renderer encodes prior observations, and adds real-data adversarial supervision to improve its visual quality. In AgentGarten, agents perceive the world through visual observations, interact with it in real time, and improve by distilling each round of experience into playbooks that subsequent agents inherit and refine. Our empirical study demonstrates a substantial gain in learning efficiency, with agents learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart. As new worlds can be written as code and rendered through the same interface, environments can scale in both number and difficulty alongside their agents, a step toward agents that keep evolving through interactive experience.

CommentsProject page: https://mirros-lab.github.io/agent-garten

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑