AI 中文总结
该研究推出面向MLE中RSI研究的开源全栈系统OpenMLE,后训练Frontis-MA1(35B)作为元进化智能体,在多个基准测试中提升性能,相关模型与系统已开源。
AI 中文摘要
递归自我改进(RSI)需要能够改进AI构建过程(即AI4AI)的AI系统,机器学习工程(MLE)为研究该能力提供了具体、可执行的测试平台。我们推出OpenMLE,这是一个面向MLE中RSI研究的开源全栈系统,涵盖带执行反馈的可验证任务环境(OpenMLE-Gym)、算子学习(OpenMLE-RL)和长程搜索(OpenMLE-Evo)。在该栈上,我们后训练Frontis-MA1(35B)作为MLE的元进化智能体,围绕四个原子程序进化算子(Draft、Improve、Debug、Crossover)对齐后训练与推理:通过基于执行的SFT和RL在针对所有评估基准去重的数据上训练相同算子,再将其组合成长程搜索,在单个循环中耦合学习与进化。在MLE-Bench Lite上,单块RTX 4090的12GB显存上限下,每任务预算12小时,Frontis-MA1(35B)结合OpenMLE-Evo将Medal Average从基础模型的39.39%提升至60.61%,结合OpenMLE-Evo-Max(基准无关经验先验与异步搜索)则达到71.21%,超过GPT-5.5 + Codex,接近GPT-5.6 Sol和2.8T Kimi K3。在保留的NatureBench Lite上,两部分均具备迁移性:框架固定时,替换为训练后的模型将Match-SOTA从50%提升至70%;模型固定时,替换为OpenMLE-Evo将其从20%提升至50%。我们发布模型权重和完整OpenMLE栈,以支持面向RSI的可执行AI4AI的可复现研究。代码:this https URL
英文摘要
Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: https://github.com/FrontisAI/OpenRSI