发表机构
Microsoft Research Montréal; Mila(微软研究院蒙特利尔分部; 米拉魁北克人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FrogNano是一个4B编码智能体,通过在线任务合成和强化学习在约1500个SWE环境中训练,证明了无需蒸馏即可用合成任务训练出有竞争力的轻量级智能体。
AI 中文摘要
我们提出了FrogNano,一个4B编码智能体,旨在高效且有效地解决软件工程(SWE)任务,即使在资源受限的环境下也是如此。它仅通过强化学习(RL)在约1,500个带有合成任务的SWE环境中进行后训练。提升性能的一个关键因素是在线任务合成流程,该流程创建的任务校准到当前检查点可学习性的前沿。本报告提供了证据表明,仅使用合成任务即可训练出具有竞争力的小型编码智能体,而无需从更大模型进行传统蒸馏,并且生成处于当前智能体可学习性前沿的任务至关重要。我们报告了训练方法、跨多种环境的评估以及深入分析的细节,这为我们持续探索可在最小硬件上运行的轻量级但功能强大的编码智能体奠定了基础。
英文摘要
We present FrogNano, a 4B coding agent designed to tackle software engineering (SWE) tasks efficiently and effectively, even under resource-constrained environments. It is post-trained exclusively via RL on around 1,500 SWE environments with synthetic tasks. A key ingredient for improving performance is an online task synthesis pipeline that creates tasks calibrated to the frontier of learnability for the current checkpoint. This report provides evidence that competitive small coding agents can be trained with synthetic tasks alone, without traditional distillation from larger models, and that generating tasks at the learnability frontier of the current agent is important. We report details on the training methodology, evaluations across diverse environments, and in-depth analyses, serving as a foundation for our ongoing exploration of lightweight yet capable coding agents that can run on minimal hardware.