arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13406cs.AIcs.LG

广义智能体迭代:迭代策略改进与递归自我改进的统一形式化框架

Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement

Hongyao Tang, Yi Ma, Pengyi Li, Yifu Yuan

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出广义智能体迭代(GAI)框架,统一描述迭代策略改进与递归自我改进,通过两个关键维度区分实例并定位现有系统,为分析设计智能体提供基础。

中文摘要 AI 辅助

当我们谈论递归自我改进(RSI)时,我们是在谈论一种现象、一种机制,还是一种前景?在迈向自主和进化智能的进程中,RSI在多个尺度上被宣称,但目前尚无一个统一框架能够形式化地描述这些新兴实例。其在经典领域的对应物——迭代策略改进——由广义策略迭代(GPI)刻画,这是一个具有广泛适用性和良好理论性质的框架,但前提是更新原则和评估基准位于智能体之外。在本文中,我们提出广义智能体迭代(GAI),这是一个形式化框架,将迭代策略改进和RSI描述为同一学习范式的两种情形。GAI将智能体定义为系统内可修改组件的配置,并将学习过程建模为智能体评估与智能体改进的循环。两个关键旋钮区分了不同实例:改进机制是否属于智能体的一部分,以及衡量标准是否锚定于智能体外部。前者划定了GPI与RSI之间的边界,后者决定了系统的极性为锚定、目标漂移或完全自指。此外,我们利用这些坐标将现有系统置于同一两个轴上,并使得递归自我改进的缺陷可以逐条件陈述。我们将本文视为探索RSI形式化表征的第一步,该表征基于经典理论,使现有系统可比较,并为分析和设计新系统提供原则性基础。

英文摘要

When we speak of recursive self-improvement (RSI), are we speaking of a phenomenon, a mechanism, or a prospect? Towards autonomous and evolving intelligence, RSI is being claimed at many scales, while no single framework that formally describes these emerging instances exists. Its counterpart in the classical realm, iterative policy improvement, is characterized by generalized policy iteration (GPI), a framework of broad applicability with well-understood theoretical properties, but only where the update principle and the evaluation base lie outside the agent. In this paper, we propose Generalized Agent Iteration (GAI), a formal framework that describes iterative policy improvement and RSI as two cases of a single learning paradigm. GAI defines the agent as a configuration of modifiable components within a system and models the learning process as a cycle of agent evaluation and agent improvement. Two pivotal dials then distinguish the instances: whether the improving mechanism is part of the agent and whether the standard it is measured against is grounded outside it. The former dial delineates the boundary between GPI and RSI, and the latter determines a system's polarity as anchored, goal drift, or fully self-referential. Moreover, we use these coordinates to place existing systems on the same two axes and make the defects of recursive self-improvement statable one condition at a time. We see this paper as a first step toward exploring a formal characterization of RSI that rests on the classical account, makes existing systems comparable, and provides a principled basis for analyzing and designing new ones.

发表机构

  • Tianjin University(天津大学)
  • Shanxi University(山西大学)

机构由 AI 辅助整理,请以论文原文为准。

↑