发表机构
University of California, Riverside (UCR); AlphaAvatar; Illinois Institute of Technology (IIT)(加州大学河滨分校; 阿尔法阿凡达; 伊利诺伊理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人工智能系统自身改进,通过对1250篇论文调查区分有界自我优化与开放式递归自我改进(RSI),考察评估器设计空间并排序信号,发现自我改进强度与层次相关,失败模式源于违规,指出治理级测量是薄弱环节。
AI 中文摘要
人工智能系统越来越多地参与自身改进,如修改输出、调整部署工具、利用自身生成的数据训练等,甚至开展人工智能研究本身。现有文献用的术语混淆了不同的目标。本文沿两个轴对1250篇arXiv论文(2024 - 2026年)进行了调查,区分了有界自我优化(已在工业实践中应用)和开放式递归自我改进(RSI)。RSI在各测量轴上受多种限制,其独特之处在于有专门的自我评估类别。文中考察了评估器设计空间,将信号按验证层次排序,观察到自我改进强度与层次相关,其失败模式源于违反层次规则,且“研究方向设定”瓶颈位于层次顶端。还将技术文献与RSI极限理论及闭环带来的安全治理问题相联系,指出自我改进的治理级测量是该领域最薄弱的环节。
英文摘要
AI systems increasingly participate in their own improvement: revising their outputs, adapting their harnesses during deployment, training on data they generate, and conducting AI research itself. This literature uses a vocabulary ("self-refine," "self-reward," "self-play," "self-evolve") that conflates fundamentally different ambitions. We survey 1,250 arXiv papers (2024-2026) along two axes: what the system improves -- its behavior in deployment, its policy through training, its evaluator, or the research process itself -- and the degree of loop closure (human-in-the-loop to fully closed). The taxonomy separates bounded self-refinement -- convergent, evaluable, and industrial practice -- from open-ended recursive self-improvement (RSI), which remains bounded by grounding requirements, collapse dynamics, and compute constraints on every measured axis. Its distinctive feature is a dedicated category for self-evaluation: every improvement loop is a claim that some signal can substitute for human judgment. We survey the evaluator design space -- judges, process reward models, verifiers, rubrics, meta-evaluation -- order the signals into a verification hierarchy from formal verifiers (strongest) to intrinsic self-assessment (weakest), and observe that demonstrated self-improvement strength tracks this hierarchy, that its failure modes (self-confirming loops, model and diversity collapse) follow from its violations, and that the "research direction-setting" bottleneck keeping humans in the loop divides into a verification problem the hierarchy indexes and a prior one -- choosing what deserves evaluation at all -- that it does not. We connect the literature to the theory of RSI limits and to the safety and governance questions raised by frontier-lab accounts of closing the loop, and identify governance-grade measurement of self-improvement as the field's most underpopulated niche.
Comments44 pages, 6 figures