arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AuraForge:扩展训练编码智能体的安全监督规模

AuraForge: Scaling Security Supervision for Training Coding Agents

Danqing Wang, Songwen Zhao, Harsh Sharma, Jierui Wang, Andre Vicente Duarte, Ivan Bercovich, Lei Li

arXiv 2610.00850首次发表:更新:

发表机构

Carnegie Mellon University; University of California, Los Angeles; ScOp Venture Capital(卡内基梅隆大学; 加州大学洛杉矶分校; ScOp风险投资公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AuraForge通过攻击导向测试合成与奖励黑客防护,构建多语言CWE训练平台AuraGym,提供更丰富可靠的安全监督,显著提升编码智能体的安全训练效果。

AI 中文摘要

编码智能体现在已足够熟练,能够从单个提示生成复杂的软件应用程序。随着其能力的增长,人工监督已越来越多地从逐行代码审查转向对结果的不干预评估。然而,近期研究表明,这种转变暴露了一个关键风险:仅凭功能正确性并不能保证实现的安全性。尽管代码安全日益受到关注,但训练更安全的编码智能体仍然具有挑战性,因为从真实世界代码库中大规模获取可靠的安全监督十分困难。我们提出AuraForge,用于合成并验证可执行的安全测试,以训练安全的编码智能体。我们的方法结合了攻击导向的测试合成、语言可扩展的任务构建以及针对奖励黑客的防护措施。利用AuraForge,我们构建了AuraGym,一个多语言、多CWE的可执行训练平台:包含来自Python、JavaScript和TypeScript的344个真实世界代码库中的679个可执行功能实现任务,覆盖177个CWE类别。在具有人工编写安全测试的子集上,AuraForge平均生成约3倍的测试用例,并将误报率降低83.23%,从而允许替代的安全实现获得正确的监督。使用合成安全测试训练Qwen3.5-4B,在三种语言上比使用人工编写安全测试获得了更大的改进(平均19.7个FuncPass和6.2个SecPass,对比14.9个FuncPass和4.4个SecPass)。这些结果表明,AuraForge为训练安全的编码智能体提供了更多样、更可靠的安全监督。

英文摘要

Coding agents are now proficient enough to generate complex software applications from a single prompt. As their capabilities have grown, human oversight has increasingly shifted from line-by-line code review toward hands-off evaluation of outcomes. However, recent studies have shown that such a transition exposes a critical risk: functional correctness alone does not guarantee a secure implementation. Despite growing attention to code security, training safer coding agents remains challenging because reliable security supervision is difficult to obtain at scale from real-world repositories. We introduce AuraForge to synthesize and validate executable security tests for training secure coding agents. Our approach combines attack-oriented test synthesis, language-extensible task construction, and safeguards against reward hacking. Using AuraForge, we construct AuraGym, a multi-language and multi-CWE executable training gym: 679 executable feature-implementation tasks from 344 real-world repositories across Python, JavaScript, and TypeScript, covering 177 CWE categories. On the subset with human-written security tests, AuraForge produces about 3 times as many test cases on average and reduces the false-positive rate by 83.23%, allowing alternative secure implementations to receive correct supervision. Training Qwen3.5-4B with synthesized security tests gains larger improvements than human-written security tests (average 19.7 FuncPass and 6.2 SecPass vs. 14.9 FuncPass and 4.4 SecPass) on three languages. These results demonstrate that AuraForge provides more diverse and reliable security supervision to train secure coding agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑