高效探索即足够
Efficient Exploration Is Enough
浏览论文内容
中文总结 AI 辅助
本文提出高效探索的新视角,证明仅通过内在目标驱动的探索即可在无外在奖励下自动产生复杂行为,为开放式学习提供原则性机制。
中文摘要 AI 辅助
本工作引入了一种关于高效探索的替代视角,并在无外在奖励的情况下研究其理论与实证意义。具体而言,我们将高效探索者定义为优先产生可泛化经验(即支持学习能够预测并适应环境的模型的数据)的智能体。这使我们能够通过预测与泛化的视角来分析高效探索。理论上,我们证明了最优高效探索者会自然地安排其轨迹,优先访问信息量最大且最易学习的区域。实证上,我们表明,优化这些智能体即使在相对简单的环境中也会产生一种逐步复杂行为的自动课程。这些结果表明,仅追求这一纯粹的内在目标就足以驱动高度复杂行为的涌现。我们认为,这一新框架提供了一种原则性机制,使智能体-环境系统能够在没有外部奖励、任务或目标的情况下,维持一个日益复杂行为的开放式过程。
英文摘要
This work introduces an alternative view of efficient exploration and studies its theoretical and empirical implications in the absence of extrinsic rewards. Specifically, we define efficient explorers as agents that prioritize generating generalizable experience, i.e., data that supports learning models capable of predicting and adapting across the environment. This allows us to analyze efficient exploration through the lens of prediction and generalization. Theoretically, we demonstrate that optimally efficient explorers naturally schedule their trajectories to visit the most informative and learnable regions first. Empirically, we show that optimizing for these agents gives rise to an automatic curriculum of progressively more complex behaviors, even in relatively simple environments. These results indicate that pursuing this purely intrinsic objective alone is enough to drive the emergence of highly sophisticated behaviors. We believe that this new framework provides a principled mechanism by which agent-environment systems may sustain an open-ended process of increasingly complex behavior without external rewards, tasks, or objectives.
发表机构
- University of the Basque Country (UPV/EHU)(巴斯克大学(UPV/EHU))
- University of Alberta(阿尔伯塔大学)
- Basque Center for Applied Mathematics (BCAM)(巴斯克应用数学中心)
机构由 AI 辅助整理,请以论文原文为准。