发表机构
Software Engineering Institute, East China Normal University; University of Chicago; Peking University(华东师范大学软件工程学院; 芝加哥大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出RAV框架,通过Route、Align、Verify三阶段协同优化,在不修改骨干模型的情况下,于MBPP基准上实现代码生成功能正确性的显著提升,性能优于基础模型。
AI 中文摘要
大语言模型(LLMs)已大幅提升代码生成能力,但要实现强功能正确性仍存在困难,尤其在异构编程任务中,单一提示策略和单一直接生成输出往往不足。本文提出RAV这一轻量模块化框架,通过三个协同阶段在固定骨干模型基础上改进代码生成:Route阶段在生成前应用任务感知提示路由;Align阶段通过对齐LoRA适配减少微调提示与推理时提示的不匹配;Verify阶段通过针对可见公开测试执行多个候选结果来选择最终输出。我们在MBPP基准的净化版和全量设置下评估RAV,完整RAV流程在所有评估配置中达到最佳性能,在MBPP净化版上得分为0.8911,MBPP全量版上得分为0.8520;与基础模型相比,这些结果分别提升了6.35和9.92个百分点。组件级消融实验进一步表明,任务感知路由和对齐适配与基于执行的验证结合时效果显著增强,额外的鲁棒性和污染分析也支持观测到的改进的可靠性。总体而言,结果表明,无需修改骨干架构,通过联合优化任务提示方式、模型适配方式及最终输出选择方式,即可显著提升代码生成的功能正确性。
英文摘要
Large language models (LLMs) have substantially improved code generation, yet achieving strong functional correctness remains difficult, especially for heterogeneous programming tasks where a single prompting strategy and a single directly generated output are often insufficient. In this paper, we present RAV, a lightweight and modular framework that improves code generation with a fixed backbone model through three coordinated stages: Route, which applies task-aware prompt routing before generation; Align, which reduces the mismatch between fine-tuning prompts and inference-time prompts through aligned LoRA adaptation; and Verify, which selects the final output by executing multiple candidates against visible public tests. We evaluate RAV on the MBPP benchmark under both the sanitized and full settings. The complete RAV pipeline achieves the best performance among all evaluated configurations, reaching 0.8911 on MBPP Sanitized and 0.8520 on MBPP Full. Compared with the base model, these results represent improvements of 6.35 and 9.92 percentage points, respectively. Component-wise ablation experiments further show that task-aware routing and aligned adaptation become substantially more effective when combined with execution-based verification. Additional robustness and contamination analyses support the reliability of the observed improvements. Overall, the results indicate that functional correctness in code generation can be meaningfully improved without modifying the backbone architecture, by jointly optimizing how tasks are prompted, how the model is adapted, and how final outputs are selected.