发表机构
SingularityNET Foundation(奇点网络基金会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文探讨内在奖励能否驱动无外部目标的适应性自组织,综述相关方法并指出失败模式,提出三个实验方向以验证内在学习能否产生更高层次的适应性组织。
AI 中文摘要
生物细胞可以被视为个体化的、相互作用的智能体,其集体动态在组织的多个层面上产生适应性行为,从单个细胞到组织,再到整个多细胞生物体。在这篇观点与教程文章中,我们讨论了人工神经系统中内在奖励是否能在没有共享外部目标的情况下支持适应性、功能特化和更高层次的自组织。我们综述了赋权、好奇心、学习进展、信息增益、无监督技能发现、互信息估计以及用于内在奖励计算的世界模型。特别关注了失败模式,展示了这些目标何时不会产生持续的探索或日益复杂的行为。我们认为,更强大的系统可能需要互补的目标、通信、记忆、多时间尺度的学习以及环境约束。基于这一观点,我们概述了三个实验方向。这些包括一个资源受限的环境,在该环境中,原本稳定的行为吸引子变得不可持续,使我们能够测试环境约束是否能缓解内在目标的典型失败模式。一个具有每智能体内在奖励的循环智能体网络,以及一个层次化世界模型智能体,在其中探索性运动能力先于目标导向行为发展。这些实验旨在测试内在学习是否能在逐步更高的层次上导致适应性组织。
英文摘要
Biological cells can be viewed as individual, interacting agents whose collective dynamics give rise to adaptive behaviour at multiple levels of organisation, from individual cells through tissues to whole multicellular organisms. In this perspective and tutorial article we discuss whether intrinsic rewards in artificial neural systems can support adaptation, functional specialisation and higher-level self-organisation without a shared external objective. We review empowerment, curiosity, learning progress, information gain, unsupervised skill discovery, mutual information estimation and the use of world models for intrinsic reward computation. Particular attention is given to failure modes showing when such objectives do not produce sustained exploration or increasingly complex behaviour. We argue that more capable systems may require complementary objectives, communication, memory, learning at multiple temporal scales and environmental constraints. Based on this perspective, we outline three experimental directions. These include a resource-constrained environment in which otherwise stable behavioural attractors become unsustainable, allowing us to test whether environmental constraints can mitigate characteristic failure modes of intrinsic objectives. The network of recurrent agents with per-agent intrinsic rewards, and a hierarchical world-model agent in which exploratory motor competence develops before goal-directed behaviour. These experiments are intended to test whether intrinsic learning can lead to adaptive organisation at progressively higher levels.