发表机构
Northeastern University; Columbia University; University of Pennsylvania; University of Southern California(东北大学; 哥伦比亚大学; 宾夕法尼亚大学; 南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出CodeForge-MA框架,通过多智能体数据锻造、执行验证强化微调和语言条件LoRA,提升多语言代码生成性能,实验显示稳健收益。
AI 中文摘要
用于代码生成的大型语言模型通常在执行、多语言覆盖和污染控制方面表现不佳,尤其是在冻结主干约束下。我们提出了CodeForge-MA,一个统一的框架,通过多智能体数据锻造、执行验证的强化指令微调和语言条件的LoRA适配器混合来改进代码合成。四个专门的智能体——Composer、Reviewer、Executor和Curator——迭代地细化指令代码对,用测试验证它们,并过滤重复和基准泄漏。在训练期间,我们将掩码监督微调与测试驱动的强化目标相结合,以使生成与可执行正确性对齐。对于更大的模型,我们使用低秩适配器上的稀疏专家路由来改善跨语言迁移,同时在推理时保持基础模型不变。实验表明,联合数据、目标和适配器设计在多种编程语言中产生了稳健的收益。
英文摘要
Large language models for code generation often fail on execution, multilingual coverage, and contamination control, especially under frozen backbone constraints. We present CodeForge-MA, a unified framework that improves code synthesis through a multi-agent data forge, execution verified reinforced instruction tuning, and a language conditioned mixture of LoRA adapters. Four specialized agents, Composer, Reviewer, Executor, and Curator, iteratively refine instruction code pairs, validate them with tests, and filter duplicates and benchmark leakage. During training, we combine masked supervised fine tuning with a test driven reinforcement objective to align generations with executable correctness. For the larger model, we use sparse expert routing over low rank adapters to improve cross language transfer while keeping the base model unchanged at inference. Experiments show that joint data, objective, and adapter design yields robust gains across programming languages.