arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25457cs.CRcs.AIcs.MA

MACGen:通过多智能体协作实现功能正确且安全的代码生成

MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration

Miseon Yu, Jaehoon Choi, Younghan Lee, Yunheung Paek

首次发表
浏览论文内容

中文总结 AI 辅助

MACGen是整合规划等多环节的多智能体代码生成框架,在CWEval、BaxBench基准上,较直接提示显著提升安全代码生成的功能与安全性指标。

中文摘要 AI 辅助

尽管大型语言模型具备强大的代码生成能力,但它们常无法生成安全代码,其输出频繁包含安全漏洞。安全代码生成本质上是一个多目标问题,需要同时满足功能正确性和安全性,这使得该任务本身极具挑战性。现有方法通过注入外部安全知识或利用智能体反馈与迭代优化来应对这一挑战,但指南检索常让生成器将通用建议转化为特定任务的安全实现,而共享对话的多智能体反馈会模糊角色边界并存在上下文膨胀问题。本文提出MACGen,这是一个整合规划、安全分析、代码合成与优化的多智能体框架,用于联合优化安全性与功能。规划器构建满足功能需求的分步计划;安全顾问识别可能的CWE(通用弱性枚举)并合成特定任务的指南;编码员基于这些构件生成代码;审核员从不同视角提供反馈。各智能体不共享完整对话历史,仅接收上游阶段的结构化构件,以此强化角色专业化并减少不受控的上下文增长。在CWEval和BaxBench基准上,MACGen的F&S@1指标较直接提示分别平均提升19.61和10.57个百分点。

英文摘要

Despite their strong ability to generate code, large language models often fail to produce secure code, as their outputs frequently contain security vulnerabilities. Secure code generation is inherently challenging because it requires solving a multi-objective problem: functional correctness and security. Existing approaches address this challenge by injecting external security knowledge or by using agentic feedback and iterative refinement. However, guideline retrieval often leaves the generator to translate generic advice into task-specific secure implementations, while shared-dialogue multi-agent feedback can blur role boundaries and suffer from context bloat. We present MACGen, a multi-agent framework that integrates planning, security analysis, code synthesis and refinement to jointly optimize security and functionality. A planner constructs a step-by-step plan to satisfy functional requirements. A security advisor identifies likely CWEs and synthesizes task-specific guidelines, a coder then generates code grounded in these artifacts, and a reviewer issues perspective-separated feedback. Rather than sharing full dialogue histories, each agent receives only structured artifacts from upstream stages, enforcing role specialization and reducing uncontrolled context growth. On CWEval and BaxBench, MACGen improves F&S@1 over direct prompting by 19.61 and 10.57 percentage points (pp) on average, respectively.

发表机构

  • Seoul National University(首尔大学)
  • Sungshin Women’s University(诚信女子大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑