arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人工实验者:利用自目的强化学习发现与控制自组织现象

The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning

Marko Cvjetko, Benedikt Hartl, Michael Levin, Clément Moulin-Frier, Pierre-Yves Oudeyer

arXiv 2608.26116首次发表:更新:

发表机构

Inria Centre at the University of Bordeaux; Allen Discovery Center at Tufts University; Wyss Institute for Biologically Inspired Engineering at Harvard University; Inria; INSA Lyon; CITI; UR3720(波尔多大学Inria中心; 塔夫茨大学艾伦发现中心; 哈佛大学Wyss生物启发工程研究所; 法国国家信息与自动化研究所; 里昂国立应用科学学院; CITI研究所; UR3720实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于自目的强化学习的闭环框架CARL,以Lenia为对象,实现了高效发现、控制自组织孤子及实时引导其穿越迷宫的能力,策略可零样本泛化至分布外条件。

AI 中文摘要

现有探索元胞自动机及其他复杂系统的方法大多以开环方式运行:它们设定初始条件,执行完整模拟,仅在结束后观察结果,执行过程中不进行干预。我们提出一种基于自目的强化学习(autotelic reinforcement learning)的闭环框架,其中智能体自主采样多样化目标,并学习目标条件策略,通过最小的局部扰动干预复杂系统。我们将该框架实例化到Lenia(一种以类生命自组织模式著称的连续元胞自动机)上,构建名为CARL的智能体系统,验证了三项能力:其一,CARL在广泛的Lenia更新规则范围内发现稳定孤子的速率高于启发式基线;其二,它仅用少量干预即可学习控制现有孤子的运动方向,表明CARL不仅能创建自组织模式,还能对其进行控制;其三,人类可通过指定高级方向指令,利用训练后的智能体实时引导孤子穿越迷宫环境,智能体将高级指令转化为低级干预。在多样化目标、更新规则及随机初始状态下训练的智能体,获得了可零样本泛化到各类分布外条件的策略。这些结果为人工实验者智能体指明了方向,这类智能体可自主或在人类引导下,发现并控制复杂系统中的涌现现象。

英文摘要

Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions, execute a full simulation, and observe the outcome, without intervening during execution. We introduce a closed-loop framework based on autotelic reinforcement learning, in which an agent autonomously samples diverse goals and learns a goal-conditioned policy to intervene in a complex system through minimal, local perturbations. We instantiate this framework on Lenia, a continuous cellular automaton known for life-like self-organizing patterns, in an agentic system we call CARL, and demonstrate three capabilities. First, CARL discovers stable solitons across a wide range of Lenia update rules at a higher rate than heuristic baselines. Second, it learns to steer the movement direction of existing solitons with few interventions, showing that CARL can control self-organizing patterns, not only create them. Third, humans can use trained agents to guide solitons through maze environments in real time by specifying high-level directional commands that the agent translates into low-level interventions. Trained across diverse goals, update rules, and random initial states, the agents acquire policies that generalize zero-shot to various out-of-distribution conditions. These results suggest a path toward artificial experimentalist agents that, autonomously or with human guidance, discover and control emergent phenomena in complex systems.

Comments8 pages, 7 figures. Accepted at ALIFE 2026. Companion website: https://developmentalsystems.org/carl/

DOI:10.1162/ISAL.a.971

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑