发表机构
Shanghai Jiao Tong University; DP Technology; National University of Singapore; Endless Frontier(上海交通大学; 深势科技; 新加坡国立大学; 无尽前沿)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究最少人类参与下,AI代理通过高密度接口和递归式自我改进,开发出在工业编码基准上领先的27B模型iCoder,展示了前沿模型开发的新路径。
AI 中文摘要
递归式AI,即AI在构建和改进AI方面扮演越来越完整角色的前景,是AI领域的一颗明珠。尽管递归式自我开发在小模型、受限任务和固定时间预算下已变得可行,但这一雄心的更具实质性的实现——即开发一个可发布、具有前沿竞争力的模型——仍然更具挑战性。在这项工作中,我们探究了代理开发一个前沿模型所需的最少人类参与程度。我们将人类输入集中于一个高密度、低频次的接口:专家将目标、阶段脚手架、权限边界和操作流程编码为可复用的研究技能,而代理则实例化这些先验知识、选择实验、诊断结果并修订训练策略。在具有挑战性的工业编码领域,该代理演化数据并协调SFT、在线策略自蒸馏以及基于可验证奖励的强化学习,最终产生了iCoder,一个用于RTL设计和GPU内核优化的27B模型。在七个基准测试中,iCoder领先于RTLLM,超越了GPT-5.5和Claude-Opus-4.8;在CVDP和KernelBench L2上排名第二,超过GPT-5.5达16个百分点;并在TritonBench上以最佳结果与Claude-Opus-4.8持平。探索性案例研究进一步表明,iCoder在竞争性迭代RTL和GPU内核优化方面表现优异,且使用的令牌数量显著更少。这些结果描绘了一条通往递归式自我改进的工程路径,其中人类提炼模型构建的原则,代理通过基于证据的实验将其付诸实践,每一代AI都成为下一代更具能力的架构师。
英文摘要
Recursive AI, the prospect of AI taking an increasingly complete role in building and improving AI, is a crown jewel of AI for AI. Although recursive self-development has become practical for small models, bounded tasks, and fixed time budgets, a more consequential realization of this ambition, i.e., developing a release-ready, frontier-competitive model, remains far more challenging. In this work, we ask how little human involvement is sufficient for an agent to develop a frontier model. We concentrate human input into a high-density, low-frequency interface: experts encode objectives, stage scaffolds, permission boundaries, and operating procedures as reusable research skills, while the agent instantiates these priors, selects experiments, diagnoses outcomes, and revises the training strategy. In the challenging domain of industrial coding, the agent evolves data and coordinates SFT, on-policy self-distillation, and reinforcement learning with verifiable rewards, ultimately producing iCoder, a 27B model for RTL design and GPU kernel optimization. Across seven benchmarks, iCoder leads RTLLM, outperforming GPT-5.5 and Claude-Opus-4.8; ranks second on CVDP and KernelBench L2, exceeding GPT-5.5 by 16 points; and ties Claude-Opus-4.8 for the best TritonBench result. Exploratory case studies further show iCoder's competitive iterative RTL and GPU-kernel optimization with substantially fewer tokens. These results chart an engineering path toward recursive self-improvement, in which humans distill the principles of model building, agents operationalize them through evidence-driven experimentation, and each generation of AI becomes a more capable architect of the next.