Metaⁿ:通过涌现深度实现递归自我改进
Meta$^n$: Recursive Self-Improvement through Emergent Depth
浏览论文内容
中文总结 AI 辅助
Metaⁿ将元操作固定后对输入递归,使各层从更高视角推理,在8个基准族上优于现有自我改进智能体,是ARC-AGI-2上唯一得分超0的模型。
中文摘要 AI 辅助
当前的自我改进大型语言模型(LLM)智能体仅优化答案,而非生成答案的过程;具备元层级的系统会将该层级固定,而自我编辑的系统为保持稳定必须保留部分编辑机制,其实现的元深度上限约为2。本文提出Metaⁿ,它将元操作Ω固定,转而对自身输入进行递归:该操作会反复应用于自身产物,读取下层求解器栈的轨迹及生成这些轨迹的代码,再将下一层写为战略预处理步骤和可调用助手库。由于Ω从不改变,它不会破坏系统稳定性;且因输入严格增长,每一层都能从比前一层更高的视角进行推理。深度由收敛而非预先设定,进化存档会在层链中搜索。在两个骨干模型上,Metaⁿ在全部8个基准族上的表现均优于现有自我改进智能体,在旨在抵制技能记忆的ARC-AGI-2任务中,它是唯一得分高于0的模型。 ablation实验表明,递归带来的增益大多来自各层向后续层传递的条件信息,且即使没有提示规定,各层也会随深度涌现出不同角色。代码可在指定URL获取。
英文摘要
Self-improving LLM agents refine answers, not the process that produces those answers. Systems that add a meta-level hold that level fixed, and those that edit themselves must leave part of their own editing machinery untouched to stay stable, capping the meta-depth they realize at roughly two. We present Meta$^n$, which keeps the meta-operation fixed and recurses on its input instead. That operation, $Ω$, is applied repeatedly to its own products, reading the traces of the solver stack below together with the code that produced them, then writing the next layer as a strategic pre-process and a library of callable helpers. Because $Ω$ never changes, it cannot destabilize the system, and because its input strictly grows, each layer reasons from a higher vantage than the last. Depth is set by convergence rather than fixed in advance, and an evolutionary archive searches over layer chains. Across two backbones, Meta$^n$ outperforms prior self-improving agents on all eight benchmark families. The sharpest case is ARC-AGI-2, built to resist skill memorization, where it alone scores above zero. Ablations indicate that most of the gain from recursion comes from the conditioning each layer passes to the next, and distinct layer roles emerge with depth although no prompt prescribes them. Code available at https://github.com/minnesotanlp/meta-n
发表机构
- University of Minnesota(明尼苏达大学)
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。