人工实验者:利用自主目标强化学习发现与控制自组织现象
文章背景与核心概要
传统的细胞自动机和复杂系统探索方法通常以“开环”方式运行:研究人员设定初始条件,运行完整的模拟,并在执行过程中不进行动态干预地观察输出。这种局限性使得系统的高效自主探索和精确控制变得困难。
本文介绍了一种由“自主目标强化学习”(autotelic reinforcement learning)驱动的闭环框架。在这种范式下,智能体自主采样不同的目标,并训练一个条件策略,利用极少量的局部扰动来影响复杂系统。作者将这种方法实例化在 Lenia(一种以生成类生命自组织模式而闻名的连续细胞自动机)中,通过一个名为 CARL 的智能体系统实现了这一突破。该研究展示了智能体在发现孤立子(Solitons)、转向控制以及实时人机协作引导复杂模式方面的强大能力,为开发能够自主或与人类协作发现和控制复杂系统中涌现现象的先进“人工实验者智能体”铺平了道路。
摘要
Traditional approaches for exploring cellular automata and complex systems typically operate in an open-loop manner: researchers establish initial conditions, run a complete simulation, and observe the output without intervening dynamically during execution.
传统的细胞自动机和复杂系统探索方法通常以“开环”方式运行:研究人员设定初始条件,运行完整的模拟,并在执行过程中不进行动态干预地观察输出。
This paper introduces a closed-loop framework powered by autotelic reinforcement learning. In this paradigm, an agent autonomously samples diverse goals and trains a goal-conditioned policy to influence complex systems using minimal, localized perturbations.
本文介绍了一种由“自主目标强化学习”(autotelic reinforcement learning)驱动的闭环框架。在这个范式中,智能体自主采样各种目标,并训练一个条件策略,通过最小的、局部的扰动来影响复杂系统。
The authors instantiate this approach in Lenia—a continuous cellular automaton celebrated for generating life-like, self-organizing patterns—via an agentic system known as CARL.
作者在 Lenia(一种以生成类生命、自组织模式而闻名的连续细胞自动机)中实例化了这种方法,通过一个被称为 CARL 的智能体系统来实现。
Key Capabilities Demonstrated
展示的核心能力
- Discovery of Solitons: CARL successfully discovers stable solitons across a wide variety of Lenia update rules at a significantly higher rate than standard heuristic baselines.
- 发现孤立子(Solitons): CARL 在各种 Lenia 更新规则中成功发现了稳定的孤立子,其发现率显著高于标准启发式基线。
- Steering and Control: The agent learns to steer the movement direction of existing solitons through minimal interventions, proving its ability to actively control self-organization patterns rather than merely generate them.
- 转向与控制: 智能体学会通过极少的干预来操控现有孤立子的运动方向,证明了其积极控制自组织模式的能力,而不仅仅是生成它们。
- Real-Time Human Guidance: Human users can leverage trained agents to guide solitons through complex maze environments in real-time, issuing high-level directional commands that the agent translates into low-level perturbations.
- 实时人类引导: 人类用户可以利用训练好的智能体,引导孤立子实时穿过复杂的迷宫环境,发出高级的方向指令,并由智能体将其转化为低级的局部扰动。
Generalization and Impact
泛化能力与影响
Trained across diverse goals, update rules, and random initial states, CARL's agents acquire robust policies that generalize zero-shot to various out-of-distribution conditions. This research paves the way for advanced artificial experimentalist agents capable of autonomously (or collaboratively with humans) discovering and controlling emergent phenomena across complex systems.
CARL 的智能体在不同的目标、更新规则和随机初始状态下进行训练,获得了强大的策略,能够零样本(zero-shot)泛化到各种分布外(out-of-distribution)条件。这项研究为先进的人工实验者智能体铺平了道路,使其能够自主(或与人类协作)发现和控制复杂系统中的涌现现象。
Article Metadata
文章元数据
- arXiv Identifier: arXiv:2608.26116 [cs.AI]
- Authors: Marko Cvjetko, Benedikt Hartl, Michael Levin, Clément Moulin-Frier, Pierre-Yves Oudeyer
- Primary Subject: Artificial Intelligence (
cs.AI) - Publication Venue: Accepted at ALIFE 2026 (8 pages, 7 figures)
- Companion Website & Demos: Developmental Systems - CARL
- License: Creative Commons Attribution 4.0

- arXiv 标识符: arXiv:2608.26116 [cs.AI]
- 作者: Marko Cvjetko, Benedikt Hartl, Michael Levin, Clément Moulin-Frier, Pierre-Yves Oudeyer
- 主要学科: 人工智能 (
cs.AI)- 发表会场: 已被 ALIFE 2026 接收(8页,7张图表)
- 配套网站与演示: Developmental Systems - CARL
- 许可协议: 知识共享署名 4.0