AI 中文总结
MetaRSI-v1提出一种元递归自改进框架,通过数据、框架和模型三类算子的计划性组合,将自改进从单一表面扩展到完整模型生产管道,并在无外部教师条件下验证于代码和科学任务。
AI 中文摘要
递归自改进(RSI)允许系统从自身的失败中改进模型构建机制,因此每个后续模型都能继承改进成果。然而,RSI几乎仅在编码和形式化基准(如科学问答和数学)上得到验证。这种格式限制将RSI局限于机器可检查的切片内进行改进,而非通用能力——在通用能力中,问题是开放的,正确性通过论证、复现或测量来确定。我们认为,RSI下一步必须在真实、多样化的科学、工程和元科学领域运作,而不仅仅是在形式化评估易于处理的地方。为此,我们提出MetaRSI-v1,其中改进是在统一范式下三个类型化算子的计划性组合。数据RSI放大现有能力并标记其边界;框架RSI在不触及权重的情况下编辑五槽脚手架;模型RSI通过受限训练将能力内化到参数中。共享一个循环内核和工件词汇表,它们使数据、脚手架和模型更改可组合而非互斥。一个双轴优化器联合决定算子顺序和每个算子的提案策略,而元级策略跨期修订计划。我们在该领域的标准评估下验证MetaRSI-v1,涉及代码和闭式科学,且无外部教师:目标模型在自身循环中扮演每个角色。MetaRSI-v1将自改进从单表面编辑重新定义为跨完整模型生产管道的组合,开辟两条路径:一条通过训练内化能力的模型路线,以及一条保持权重不变从而将自改进扩展到任何可通过接口访问的模型的框架路线,其中数据RSI被重新定义为滋养两者的共享基底。该框架进一步产生关于循环存在于何处、算子如何组合以及监督带来什么的可反驳规律。
英文摘要
Recursive self-improvement (RSI) lets a system improve the model-building machinery from its own failures, so every later model inherits the gain. Yet RSI has been validated almost exclusively on coding and formal benchmarks such as science QA and mathematics. This format bound limits RSI to improvement within a machine-checkable slice, not general capability where questions are open and correctness is settled by argument, replication, or measurement. We argue RSI must next operate across real, diverse scientific, engineering, and meta-scientific domains, not where formal evaluation is merely tractable. To that end we present MetaRSI-v1, where improvement is the scheduled composition of three typed operators over one unified paradigm. Data-RSI amplifies existing competence and marks its boundary; Harness-RSI edits a five-slot scaffold without touching weights; Model-RSI internalizes capability into parameters through bounded training. Sharing one loop kernel and artifact vocabulary, they make data, scaffold, and model changes composable rather than exclusive. A two-axis optimizer jointly decides operator order and each operator's proposal policy, while a meta-level policy revises the schedule across terms. We validate MetaRSI-v1 under the field's standard evaluations, on code and closed-form science, with no external teacher: the target model plays every role in its own loop. MetaRSI-v1 reframes self-improvement from a single-surface edit to a composition across the full model-production pipeline, opening two paths: a model route internalizing capability through training, and a harness route leaving weights untouched and thus extending self-improvement to any model reachable through an interface, with Data-RSI redefined as the shared substrate feeding both. The framework further yields refutable laws on where loops exist, how operators compose, and what supervision buys.
Comments47 pages, 12 figures, 11 tables