arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MindForge:通过无源码程序合成教小型语言模型全生命周期软件工程

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

Yihao Chen, Shi Chang, Khaled Chawa, Feng Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan

arXiv 2607.27146首次发表:更新:

发表机构

Huawei Canada; Queen’s University; University of Manitoba(华为加拿大; 女王大学; 曼尼托巴大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对小型语言模型从零构建程序的挑战,提出MindForge构建全生命周期无源码训练环境,微调Qwen3.6-27B后在ProgramBench及7个软件工程基准上性能显著提升。

AI 中文摘要

编码智能体在修改现有代码库的软件工程任务(包括漏洞修复和功能实现)上已取得显著进展,但从零构建完整程序仍是重大挑战:即使在ProgramBench上评估的前沿模型,也仅能完全解决不到1%的任务。一个障碍是从零构建场景缺乏可扩展的训练环境,覆盖整个软件工程生命周期,而现有环境构建框架仅关注软件开发的单一阶段。为解决这一缺口,我们提出MindForge,这是一种自动化流水线,可将开源命令行程序转换为仅暴露编译后的参考可执行文件及其文档的无源码环境。利用MindForge,我们从与ProgramBench不重叠的代码库构建训练环境,并整理高质量数据配方,包含以GLM-5.2作为教师智能体的程序合成轨迹。在这些轨迹上微调Qwen3.6-27B,使其在ProgramBench上的平均测试通过率从37.98%提升至49.51%,达到与大得多的前沿模型相当的性能。此外,微调后的模型在全部7个未见过的软件工程基准上均持续优于基础模型,涵盖长周期代码库生成与转换、漏洞修复、功能实现及跨语言问题解决,在RepoZero-C2Rust上绝对提升31.00个百分点,在DeepSWE上提升14.16个百分点,在NL2Repo-Bench(含/不含测试)上分别提升10.70/4.56个百分点,在SWE-bench Verified上提升5.04个百分点,在SWE-bench Pro上提升5.93个百分点,在SWE-bench Multilingual上提升5.22个百分点,在FeatBench上提升4.94个百分点。

英文摘要

Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolve fewer than 1% of tasks. One obstacle is the lack of scalable training environments for this from-scratch setting, spanning the whole software engineering life cycle, as existing environment-construction frameworks focus only on a single phase in software development. To address this gap, we introduce MindForge, an automated pipeline that converts open-source command-line programs into source-free environments that expose only a compiled reference executable and its documentation. Using MindForge, we construct training environments from repositories disjoint from those in ProgramBench, and curate a high-quality data recipe consisting of program synthesis trajectories using GLM-5.2 as the teacher agent. Fine-tuning Qwen3.6-27B on these trajectories increases its ProgramBench average test pass rate from 37.98% to 49.51%, achieving performance comparable to substantially larger frontier models. Moreover, the fine-tuned model consistently improves over the base model across all seven unseen software engineering benchmarks, spanning long-horizon repository generation and translation, bug fixing, feature implementation, and cross-language issue resolution, with absolute gains of 31.00 points on RepoZero-C2Rust, 14.16 on DeepSWE, 10.70/4.56 on NL2Repo-Bench (with/without tests), 5.04 on SWE-bench Verified, 5.93 on SWE-bench Pro, 5.22 on SWE-bench Multilingual, and 4.94 on FeatBench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑