arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Frontis-MA1:训练AI4AI模型以实现机器学习工程中的递归自我改进

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Junlin Yang, Che Jiang, Yu Fu, Tianwei Luo, Can Ren, Weizhi Wang, Kaikai Zhao, Hongyi Liu, Yuxin Zuo, Yuru Wang, Yuchen Fan, Kai Tian, Zhenzhao Yuan, Xiaojian Lin, Li Sheng, Rushi Qiang, Guoli Jia, Xingtai Lv, Ermo Hua, Dianqiao Lei, Youbang Sun, Ning Ding, Bowen Zhou, Kaiyan Zhang

arXiv 2607.28568首次发表:更新:

AI 中文总结

该研究推出面向MLE中RSI研究的开源全栈系统OpenMLE,后训练Frontis-MA1(35B)作为元进化智能体,在多个基准测试中提升性能,相关模型与系统已开源。

AI 中文摘要

递归自我改进(RSI)需要能够改进AI构建过程(即AI4AI)的AI系统,机器学习工程(MLE)为研究该能力提供了具体、可执行的测试平台。我们推出OpenMLE,这是一个面向MLE中RSI研究的开源全栈系统,涵盖带执行反馈的可验证任务环境(OpenMLE-Gym)、算子学习(OpenMLE-RL)和长程搜索(OpenMLE-Evo)。在该栈上,我们后训练Frontis-MA1(35B)作为MLE的元进化智能体,围绕四个原子程序进化算子(Draft、Improve、Debug、Crossover)对齐后训练与推理:通过基于执行的SFT和RL在针对所有评估基准去重的数据上训练相同算子,再将其组合成长程搜索,在单个循环中耦合学习与进化。在MLE-Bench Lite上,单块RTX 4090的12GB显存上限下,每任务预算12小时,Frontis-MA1(35B)结合OpenMLE-Evo将Medal Average从基础模型的39.39%提升至60.61%,结合OpenMLE-Evo-Max(基准无关经验先验与异步搜索)则达到71.21%,超过GPT-5.5 + Codex,接近GPT-5.6 Sol和2.8T Kimi K3。在保留的NatureBench Lite上,两部分均具备迁移性:框架固定时,替换为训练后的模型将Match-SOTA从50%提升至70%;模型固定时,替换为OpenMLE-Evo将其从20%提升至50%。我们发布模型权重和完整OpenMLE栈,以支持面向RSI的可执行AI4AI的可复现研究。代码:this https URL

英文摘要

Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: https://github.com/FrontisAI/OpenRSI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑