ExecuGraph:一种用于通过大语言模型进行可靠后端代码合成的多智能体、基于执行的框架
ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models
浏览论文内容
中文总结 AI 辅助
研究针对大语言模型后端代码合成可靠性问题,提出ExecuGraph多智能体框架,将执行验证置于核心,通过六个智能体和有向工作流协调,在多数据集评估中对比单智能体基线,验证方法有效性且能控制测量各因素贡献。
中文摘要 AI 辅助
大语言模型能生成看似合理的后端代码,但单步范式无法保证正确性或运行时可靠性。我们提出了ExecuGraph,这是一个多智能体框架,将基于执行的验证置于后端代码合成的核心。六个专门的智能体(规划器、代码生成器、逻辑审查器、评估器、优化器和解释器)由一个具有有限重试预算的类型化有向工作流协调,在LangGraph上使用本地托管模型(Ollama)实现,并带有一个用于算法技术召回的可选检索层。一个带有挂钟超时的子进程隔离沙盒保护每次评估。我们在精心策划的30个问题的DSA套件(internal - 30)、HumanEval(n = 64)和一个APPS入门子集中进行评估,将ExecuGraph与单智能体一次性基线和单智能体执行重试基线(一种隔离多智能体分解贡献的Reflexion风格消融)进行对比。在internal - 30上,三种情况在统计上无显著差异;在HumanEval上,多智能体方法领先3.1个百分点。最强的信号是跨模型的:使用DeepSeekCoder V2 Lite时,图类别准确率从57.5%(一次性)提高到80.0%(多智能体完整)。该框架的主要贡献是方法学上的:一个单一代码库可通过配置分解为一次性、执行重试和每个智能体的消融条件,从而能够控制测量每个因素的边际贡献。还报告了每个智能体的消融、重试预算扫描、错误类别分类和测试源审核。
英文摘要
Large Language Models generate plausible backend code, but a single-pass paradigm provides no guarantee of correctness or runtime reliability. We present ExecuGraph, a multi-agent framework that places execution-based validation at the center of backend code synthesis. Six specialized agents (Planner, Code Generator, Logical Reviewer, Evaluator, Optimizer, and Explainer) are coordinated by a typed directed workflow with a bounded retry budget, implemented on LangGraph with locally hosted models (Ollama) and an optional retrieval layer for algorithmic technique recall. A subprocess-isolated sandbox with a wall-clock timeout guards every evaluation. We evaluate on a curated 30-problem DSA suite (internal-30), HumanEval (n=64), and an APPS-introductory subset, contrasting ExecuGraph against a single-agent one-shot baseline and a single-agent execution-retry baseline (a Reflexion-style ablation that isolates the contribution of multi-agent decomposition). On internal-30, the three conditions are statistically indistinguishable (n=30; paired Wilcoxon p=0.59 MF vs. SO, p=0.08 SR vs. SO); 95% bootstrap confidence intervals on all pairwise mean differences include zero. On HumanEval, multi-full edges ahead by +3.1 pp. The strongest signal is cross-model: with DeepSeekCoder V2 Lite, graph-category accuracy improves from 57.5% (oneshot) to 80.0% (multi-full), a +22.5 pp jump that supports a scaling hypothesis: the value of multi-agent decomposition grows with base-model capability. The framework's primary contribution is methodological: a single codebase that collapses by configuration into one-shot, execution-retry, and per-agent ablation conditions, enabling controlled measurement of each lever's marginal contribution. A per-agent ablation, retry-budget sweep, error-class taxonomy, and test-source audit are reported.
发表机构
- Department of Computer Science and Engineering (AI & ML), Kakatiya Institute of Technology and Science(卡凯蒂亚理工学院计算机科学与工程系(人工智能与机器学习))
- Department of Computer Science and Engineering, Kakatiya Institute of Technology and Science(卡凯蒂亚理工学院计算机科学与工程系)
- Boston University(波士顿大学)
机构由 AI 辅助整理,请以论文原文为准。