arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11394cs.SE

GraphAlignCoder:对齐程序与证明图的代码生成框架

GraphAlignCoder: Aligning Program and Proof Graphs for Code Generation

Yueke Zhang, Zihan Fang, Kevin Leach, Yu Huang

首次发表
浏览论文内容

中文总结 AI 辅助

GraphAlignCoder是将显式正确性结构迁移至代码生成的训练框架,通过对齐程序与证明图提升性能,在多个基准测试中优于现有方法,验证图注入与验证到代码的整合是关键。

中文摘要 AI 辅助

代码大语言模型(LLMs)可生成语法合理但违反隐含语义约束的程序。现有基于执行反馈的训练方法仅能判断完成的程序是否失败,却无法提供关于正确解决方案应如何组织的充足监督。我们提出GraphAlignCoder,这是一种将显式正确性结构迁移至代码生成的训练框架。GraphAlignCoder构建实现图以捕获程序区域间的控制与依赖关系;并行地,受约束的Lean流水线生成证明轨迹,从中提取形式化证明流图。模型首先学习可执行代码及图衍生的各程序区域正确性依据描述,再将该知识整合至代码生成中。GraphAlignCoder在所有基准测试中均优于基础模型、仅代码的SFT和CodeRL:与CodeRL相比,其在LiveCodeBench v6上的已解决任务数从38增至50(相对提升31.6%),在BigCodeBench Hard上从16增至23(相对提升43.8%),同时将BigCodeBench Full的任务数从359提升至363。消融研究进一步表明,验证图注入是初始推理增益的来源,而验证到代码的整合对实现稳健的跨基准测试迁移至关重要。

英文摘要

Code large language models (LLMs) can generate syntactically plausible programs that nevertheless violate hidden semantic constraints. Existing execution-feedback training methods identify whether a completed program fails, but provide limited supervision about how a correct solution should be organized. We introduce GraphAlignCoder, a training framework that transfers explicit correctness structure into code generation. GraphAlignCoder constructs an implementation graph that captures control and dependence among program regions. In parallel, a constrained Lean pipeline produces proof traces, from which we extract a formal proof-flow graph. The model first learns executable code together with graph-derived descriptions of why individual program regions are correct, and then consolidates this knowledge into code generation. GraphAlignCoder consistently outperforms the base model, code-only SFT, and CodeRL across all benchmarks. Compared with CodeRL, it increases the solved count from 38 to 50 on LiveCodeBench v6 and from 16 to 23 on BigCodeBench Hard, corresponding to relative gains of 31.6% and 43.8%, while also improving BigCodeBench Full from 359 to 363 tasks. The ablation study further shows that verification-graph injection produces the initial reasoning gain, while verification to code consolidation is essential for robust cross-benchmark transfer.

补充信息

↑